loki-mode 9.8.0 → 9.12.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (79) hide show
  1. package/README.md +19 -14
  2. package/SKILL.md +3 -2
  3. package/VERSION +1 -1
  4. package/autonomy/loki +122 -1
  5. package/autonomy/run.sh +49 -2
  6. package/dashboard/__init__.py +1 -1
  7. package/dashboard/api_evidence.py +411 -0
  8. package/dashboard/api_operator.py +283 -0
  9. package/dashboard/api_phases.py +262 -0
  10. package/dashboard/api_releases.py +242 -0
  11. package/dashboard/api_runs.py +477 -0
  12. package/dashboard/api_tests.py +444 -0
  13. package/dashboard/api_v2.py +47 -1
  14. package/dashboard/server.py +54 -0
  15. package/dashboard/static/index.html +246 -135
  16. package/docs/ARCHITECTURE-OVERVIEW.md +5 -3
  17. package/docs/CAPABILITY-BACKLOG.md +53 -0
  18. package/docs/COMPARISON.md +2 -2
  19. package/docs/COMPETITIVE-ANALYSIS.md +1 -1
  20. package/docs/COMPETITIVE-SCORECARD.md +422 -0
  21. package/docs/DASHBOARD-9.12-EVIDENCE.md +97 -0
  22. package/docs/DASHBOARD-ARCHITECTURE.md +423 -0
  23. package/docs/DEMOS.md +21 -23
  24. package/docs/HANDOFF-2026-08-03.md +439 -0
  25. package/docs/INSTALLATION.md +17 -10
  26. package/docs/OUTCOME-FRONTIER.md +536 -0
  27. package/docs/PROMPT-ABLATION-RESULT.md +97 -0
  28. package/docs/TOOLS.md +800 -0
  29. package/docs/alternative-installations.md +2 -3
  30. package/docs/audit-logging.md +44 -35
  31. package/docs/authentication.md +13 -2
  32. package/docs/authorization.md +87 -81
  33. package/docs/git-workflow.md +6 -3
  34. package/docs/metrics.md +15 -16
  35. package/docs/network-security.md +16 -13
  36. package/docs/openclaw-integration.md +36 -556
  37. package/docs/show-hn-post.md +2 -2
  38. package/docs/siem-integration.md +39 -36
  39. package/loki-ts/dist/loki.js +18 -18
  40. package/mcp/__init__.py +1 -1
  41. package/package.json +2 -2
  42. package/plugins/loki-mode/.claude-plugin/plugin.json +1 -1
  43. package/references/confidence-routing.md +18 -1
  44. package/references/invariant-checks.md +13 -8
  45. package/references/magic-rarv-integration.md +0 -1
  46. package/references/multi-provider.md +27 -5
  47. package/skills/healing.md +4 -2
  48. package/tools/audit-docs.py +488 -0
  49. package/tools/baseline-pin.py +19 -1
  50. package/tools/calibration-audit.py +523 -0
  51. package/tools/ci-gate.py +19 -1
  52. package/tools/cost-forecast.py +344 -0
  53. package/tools/cost-guard.py +19 -1
  54. package/tools/cost-history.py +19 -1
  55. package/tools/cost-per-outcome.py +394 -0
  56. package/tools/estimate-run.py +19 -1
  57. package/tools/evidence-freshness.py +307 -0
  58. package/tools/gate-init.py +19 -1
  59. package/tools/gate-report.py +19 -1
  60. package/tools/gate-simulate.py +570 -0
  61. package/tools/gate-trend.py +354 -0
  62. package/tools/model-advisor.py +52 -1
  63. package/tools/policy-load.py +19 -1
  64. package/tools/prompt-cost.py +363 -0
  65. package/tools/prompt-diff.py +448 -0
  66. package/tools/prompt-lint.py +448 -0
  67. package/tools/receipt-bundle.py +72 -2
  68. package/tools/receipt-diff.py +19 -1
  69. package/tools/receipt-find.py +19 -1
  70. package/tools/receipt-stats.py +380 -0
  71. package/tools/receipt-timeline.py +478 -0
  72. package/tools/receipt-verify-batch.py +291 -0
  73. package/tools/run-replay.py +19 -1
  74. package/tools/signing-status.py +19 -1
  75. package/tools/token-guard.py +19 -1
  76. package/tools/token-tax.py +375 -0
  77. package/tools/tool-index.py +19 -1
  78. package/tools/verification-tax.py +277 -0
  79. package/tools/verify-chain.py +361 -0
package/docs/TOOLS.md ADDED
@@ -0,0 +1,800 @@
1
+ # Bundled tools
2
+
3
+ `tools/` holds 32 standalone scripts that ship with loki-mode. They read
4
+ artifacts a run leaves behind and answer one question each. None of them starts
5
+ a build or spends money on inference.
6
+
7
+ Two qualifications to that, because "read-only" is not true of all 32:
8
+ `probe-model-catalog.py` makes network requests to provider documentation pages,
9
+ and several write files (`gate-init.py`, `cost-history.py record`,
10
+ `baseline-pin.py set`, `receipt-export.py --out`, `gate-log.py record`). Where a
11
+ tool asserts read-only in its own help, this page repeats the assertion;
12
+ `preflight.sh`, `gate-status.py`, `cost-attribute.py` and `run-replay.py` do.
13
+
14
+ Every command and every output on this page was executed against loki-mode
15
+ 9.8.1 before being written down. Where a tool reports UNKNOWN or refuses to
16
+ answer, that real output is shown rather than a prettier invented one.
17
+
18
+ ## Running them
19
+
20
+ Most of these files are not marked executable, so invoke them through the
21
+ interpreter rather than as `./tools/x.py`:
22
+
23
+ ```
24
+ python3 tools/<name>.py [args]
25
+ bash tools/<name>.sh [args]
26
+ ```
27
+
28
+ Installed from npm, they live under the package root. `python3 tools/tool-index.py`
29
+ lists what shipped.
30
+
31
+ **Where you run them from matters.** The receipt and gate tools check drift and
32
+ the working tree against the **current directory**, and several read or write
33
+ their state files relative to cwd (`.loki-policy.json`, `.loki/cost-history.jsonl`,
34
+ `.loki/gate-log.jsonl`, `.loki/baseline.json`). Run those from the workspace the
35
+ receipt describes, not from the loki-mode package directory. The examples below
36
+ show this as an explicit `cd`, with the tool referenced through a `$LOKI`
37
+ variable pointing at the package root:
38
+
39
+ ```
40
+ LOKI=/path/to/loki-mode # or the npm package root
41
+ cd /path/to/your/workspace
42
+ python3 $LOKI/tools/receipt-attest.py .loki/proofs/good/proof.json
43
+ ```
44
+
45
+ Tools that take a workspace as a positional argument and read nothing relative
46
+ to cwd (`cost-attribute.py`, `run-replay.py`, `receipt-find.py`,
47
+ `estimate-run.py`, `model-advisor.py`, `preflight.sh`) can be run from anywhere.
48
+
49
+ ## The exit-code convention
50
+
51
+ The tools that compose in CI share one convention. It is enforced by
52
+ `tests/test_tool_exit_contract.py`, which walks the real directory rather than
53
+ testing tools one at a time.
54
+
55
+ Enforcement covers exactly **24** of them: that test globs `tools/*.py` (30
56
+ files) and excludes six by name. The two shell tools, `preflight.sh` and
57
+ `verify-demo.sh`, are never globbed by it at all; they follow the convention by
58
+ observation here, not by enforcement.
59
+
60
+ | Code | Meaning |
61
+ |---|---|
62
+ | 0 | Checked, and it passed |
63
+ | 1 | Checked, and it FAILED |
64
+ | 2 | Could NOT be checked; no policy configured, or no measurement exists |
65
+ | 3 | Nothing to check (empty) |
66
+ | 64 | Usage error |
67
+ | 66 | Input missing (the path does not exist) |
68
+
69
+ Two properties are worth knowing before you build a gate on these.
70
+
71
+ **2 is not a pass.** A gate that cannot evaluate its axis exits 2, never 0. If
72
+ cost was never measured, `cost-guard.py` refuses to say the run was within
73
+ budget, because it does not know. Treat any code `>= 2` as "do not merge on
74
+ this axis".
75
+
76
+ **Gates and advisors differ deliberately.** A tool named `*-gate`, `*-guard`,
77
+ or `gate-*`, plus `ci-gate.py`, `receipt-attest.py` and `receipt-bundle.py`, is
78
+ a gate: its exit code is a verdict a machine branches on, and it may never exit
79
+ 0 while its own output says it could not evaluate. An advisor such as
80
+ `model-advisor.py` answering "NO BASIS" has answered the question honestly and
81
+ exits 0. Forcing advisors non-zero would train operators to ignore them.
82
+
83
+ Six Python tools are excluded from the convention by name, and are described
84
+ under [Internal and maintenance](#internal-and-maintenance) below. They are
85
+ benchmark harnesses and repo-maintenance scripts that predate it. One of the six,
86
+ `hybrid_search.py`, is genuinely user-facing and is documented under
87
+ [Diagnose a run](#diagnose-a-run) instead.
88
+
89
+ For the exit codes of the `loki` CLI itself (`loki start`, `loki verify`,
90
+ `loki ci`, `loki doctor`), see [exit-codes.md](./exit-codes.md). That is a
91
+ separate contract; this page does not restate it.
92
+
93
+ ---
94
+
95
+ ## Verify a receipt
96
+
97
+ An Evidence Receipt (`proof.json`) is what a run leaves behind to describe what
98
+ it did. These tools re-check one.
99
+
100
+ ### `verify-demo.sh`
101
+
102
+ **Answers:** what does the verification chain actually do, without paying for a
103
+ build? It creates a scratch git repo in a tempdir, generates a real receipt over
104
+ it, verifies it, then tampers with the receipt and verifies again so you can see
105
+ the failure. Synthetic receipts, real verifier: every verdict shown is printed by
106
+ the actual tool.
107
+
108
+ ```
109
+ $ bash tools/verify-demo.sh
110
+ ...
111
+ STEP 3 of 3 -- tamper with the receipt, then re-verify
112
+ edited facts.git.diff.count: 1 -> 999 (claims work that is not there)
113
+ $ tools/receipt-attest.py .loki/proofs/tampered/proof.json
114
+
115
+ FAILED -- drift, integrity did not pass here; receipt is UNSIGNED, so origin rests on its generator
116
+
117
+ integrity FAILED
118
+ hash mismatch: recorded 34175ce3..., computed 64eb7c52... -- proof.json was edited after it was written
119
+
120
+ -> exit 1. Caught, with the reason printed above.
121
+ ```
122
+
123
+ Pass `--keep` to leave the scratch directory in place. Exits 0 when the whole
124
+ chain behaves as expected, so it doubles as a smoke test.
125
+
126
+ ### `receipt-attest.py`
127
+
128
+ **Answers:** does this receipt hold up, checked from here? Scores five axes
129
+ independently (integrity, drift, headline, cost, tree) plus signature state, and
130
+ never collapses "could not check" into either pass or fail.
131
+
132
+ ```
133
+ $ cd /path/to/your/workspace
134
+ $ python3 $LOKI/tools/receipt-attest.py .loki/proofs/good/proof.json
135
+ VERIFIED -- every axis checked here passed; receipt is UNSIGNED, so origin rests on its generator
136
+
137
+ receipt proof.json
138
+ sha256 5854c7a5271d567f915e495b81644551f7418d485422ba1e9dcc3f62f364edc8
139
+ checked from /private/var/folders/.../loki-doccheck2-pdtkmezm
140
+
141
+ integrity VERIFIED
142
+ drift VERIFIED
143
+ headline VERIFIED
144
+ cost VERIFIED
145
+ tree VERIFIED
146
+ signature UNSIGNED (unsigned)
147
+ the receipt carries no gpg signature, so its origin rests on the generator that produced it, not on cryptographic proof
148
+ ```
149
+ (the sha256 and path are specific to that receipt; yours will differ)
150
+
151
+ Exit 0 all axes verified, 1 something FAILED, 2 something was UNVERIFIABLE, 64
152
+ usage error.
153
+
154
+ **Important:** the drift and tree axes are checked against the **current working
155
+ directory**, and this tool has no flag to point them elsewhere. Run it from the
156
+ repository the receipt describes. Run it from anywhere else and you get a
157
+ truthful but useless FAILED, because the tree genuinely does not match:
158
+
159
+ ```
160
+ $ cd /path/to/loki-mode # the WRONG directory for this receipt
161
+ $ python3 tools/receipt-attest.py /tmp/other-repo/.loki/proofs/good/proof.json
162
+ FAILED -- drift, tree did not pass here
163
+ drift FAILED
164
+ diff drift: the receipt recorded 1 files / +2 / -0, the repository now has 4178 files / +878921 / -0
165
+ ```
166
+
167
+ `receipt-bundle.py` and `receipt-export.py` accept `--repo-dir` for exactly this
168
+ case; `receipt-attest.py` does not.
169
+
170
+ ### `receipt-bundle.py`
171
+
172
+ **Answers:** do ALL the receipts under this workspace hold up, as one audit
173
+ trail? A bundle is only as good as its weakest receipt, and failures are counted,
174
+ never dropped.
175
+
176
+ ```
177
+ $ python3 tools/receipt-bundle.py /tmp/ws
178
+ Receipt bundle: /tmp/ws
179
+
180
+ VERIFIED /tmp/ws/.loki/proofs/good/proof.json
181
+
182
+ total cost UNKNOWN (0 of 1 receipts measured cost)
183
+
184
+ VERIFIED -- all 1 receipts verified. total cost UNKNOWN (0 of 1 receipts measured cost)
185
+ ```
186
+
187
+ Takes `--repo-dir` to name the repository the receipts are re-checked against.
188
+
189
+ ### `receipt-export.py`
190
+
191
+ **Answers:** can I hand a third party one self-describing evidence file? Writes
192
+ every receipt under a workspace into a single record, and includes an explicit
193
+ "what this export does NOT prove" section rather than letting the reader assume.
194
+
195
+ ```
196
+ $ python3 tools/receipt-export.py /tmp/ws --out /tmp/ws/evidence.json
197
+ ...
198
+ What this export does NOT prove:
199
+ - An UNSIGNED receipt proves INTEGRITY, not ORIGIN: the recorded bytes were not
200
+ edited after they were hashed, but nothing here shows WHICH machine produced them.
201
+ ```
202
+
203
+ `--force` overwrites an existing `--out`. Exits 1 when any receipt fails.
204
+
205
+ ### `receipt-diff.py`
206
+
207
+ **Answers:** what actually changed between two runs? Reports UNKNOWN where a
208
+ value was never measured, rather than printing a misleading zero.
209
+
210
+ ```
211
+ $ python3 tools/receipt-diff.py a/proof.json b/proof.json
212
+ Evidence Receipt diff
213
+ cost UNKNOWN -> UNKNOWN delta UNKNOWN
214
+ cache hit ratio UNKNOWN -> UNKNOWN delta UNKNOWN
215
+ iterations 0 -> 0 delta +0
216
+ duration 0 -> 0 delta +0s
217
+
218
+ gates no verdict changes
219
+
220
+ NOT COMPARABLE (reported UNKNOWN, not zero):
221
+ cost: cost was not measured in a/proof.json and b/proof.json
222
+ ```
223
+
224
+ It refuses to diff a receipt that fails integrity, rather than comparing numbers
225
+ that were edited after the fact:
226
+
227
+ ```
228
+ $ python3 tools/receipt-diff.py good/proof.json tampered/proof.json
229
+ REFUSED: receipt failed integrity verification: tampered/proof.json
230
+ hash mismatch: recorded 34175ce3..., computed 64eb7c52... -- proof.json was edited after it was written
231
+ ```
232
+ Exit 2, because nothing could be checked.
233
+
234
+ ### `receipt-find.py`
235
+
236
+ **Answers:** which receipts, out of a workspace full of them, does an auditor
237
+ actually need? Filters by measurable criteria only.
238
+
239
+ ```
240
+ $ python3 tools/receipt-find.py /tmp/ws
241
+ /tmp/ws/.loki/proofs/good/proof.json [-]
242
+ /tmp/ws/.loki/proofs/tampered/proof.json [-]
243
+
244
+ 2 of 2 receipt(s) matched. Filters applied: none (no filter applied).
245
+ ```
246
+
247
+ Filters: `--min-usd`, `--max-usd`, `--failed-only`, `--since YYYY-MM-DD`.
248
+
249
+ ### `signing-status.py`
250
+
251
+ **Answers:** can this machine produce SIGNED receipts? It proves the answer by
252
+ attempting a sign-and-verify round trip rather than assuming from config.
253
+
254
+ ```
255
+ $ python3 tools/signing-status.py
256
+ Receipt signing: UNSIGNED receipts prove integrity but NOT origin
257
+
258
+ gpg installed: /opt/homebrew/bin/gpg
259
+ LOKI_PROOF_GPG_KEY: not set
260
+ sign+verify proof: not proven
261
+
262
+ Why: LOKI_PROOF_GPG_KEY is not set
263
+
264
+ Nothing is broken. Signing is opt-in and off. Turn it on to prove
265
+ a receipt came from you and not merely that its bytes are intact.
266
+
267
+ Next: gpg --list-secret-keys --keyid-format=long # then: export LOKI_PROOF_GPG_KEY=<key-id>
268
+ ```
269
+ Exit 2 here: signing is off, so the question could not be answered affirmatively.
270
+
271
+ ---
272
+
273
+ ## Set up a merge gate
274
+
275
+ One exit code a CI job branches on, plus the tools that configure, render, and
276
+ track it.
277
+
278
+ ### `gate-status.py`
279
+
280
+ **Answers:** is this repo's merge gate actually set up and working? Start here.
281
+ One screen, read-only.
282
+
283
+ ```
284
+ $ python3 tools/gate-status.py /tmp/ws
285
+ Merge gate status for /tmp/ws
286
+ (read-only: starts nothing, spends nothing, contacts no provider)
287
+
288
+ COMPONENT STATE DETAIL
289
+ policy OK /tmp/ws/.loki-policy.json is valid and enforces: --require-receipt
290
+ baseline PROBLEM NO BASELINE: no baseline pinned. Pin one with: tools/baseline-pin.py set <workspace>
291
+ signing UNKNOWN signing is opt-in and off (LOKI_PROOF_GPG_KEY unset)
292
+ cost_history UNKNOWN no measured cost history; record a run first.
293
+ would_run OK yes: ci-gate would enforce the loaded policy on the next run
294
+
295
+ GATE: UNKNOWN -- the merge gate cannot be fully verified from here, so it is not known to be working
296
+ ```
297
+ Exit 2: a gate that is not known to be working is not reported as working.
298
+
299
+ ### `gate-init.py`
300
+
301
+ **Answers:** what policy file and CI snippet turn the gate on? Scaffolds them
302
+ from measured history, and refuses to invent a cost ceiling it has no basis for.
303
+
304
+ ```
305
+ $ cd /path/to/your/workspace
306
+ $ python3 $LOKI/tools/gate-init.py .
307
+ wrote .loki-policy.json
308
+ {
309
+ "require_receipt": true
310
+ }
311
+
312
+ validated by policy-load.py; ci-gate args: --require-receipt
313
+
314
+ NO CEILING WAS SET: no measured runs in this workspace's cost history.
315
+ The max_usd key is ABSENT rather than guessed. You must choose a ceiling and add it, for example:
316
+ "max_usd": 5.00
317
+ ```
318
+ `--print-workflow` also prints a starting-point CI snippet. `--force` overwrites.
319
+
320
+ ### `ci-gate.py`
321
+
322
+ **Answers:** does this run pass every configured merge policy? This is the one a
323
+ CI job calls.
324
+
325
+ ```
326
+ $ cd /path/to/your/workspace
327
+ $ python3 $LOKI/tools/ci-gate.py . --require-receipt
328
+ POLICY STATE DETAIL
329
+ receipt PASS VERIFIED
330
+
331
+ GATE: PASS -- 1 of 1 policies passed
332
+ ```
333
+
334
+ And on a workspace whose receipt does not hold up:
335
+
336
+ ```
337
+ $ python3 $LOKI/tools/ci-gate.py . --require-receipt
338
+ POLICY STATE DETAIL
339
+ receipt FAIL FAILED
340
+
341
+ GATE: FAIL -- 0 of 1 policies passed
342
+ ```
343
+ Exit 1. `--max-usd N` adds the cost ceiling. `--json` emits the verdict that the
344
+ four renderers below consume on stdin.
345
+
346
+ The `--require-receipt` policy runs `receipt-attest.py` underneath, so it
347
+ inherits the cwd dependence described above. Run it from the workspace; run it
348
+ from elsewhere and it FAILs on drift even for a sound receipt.
349
+
350
+ ### `policy-load.py`
351
+
352
+ **Answers:** what does my version-controlled policy file actually enforce? Keeps
353
+ the policy reviewable in git instead of buried in a CI flag.
354
+
355
+ ```
356
+ $ cd /path/to/your/workspace # where .loki-policy.json lives
357
+ $ python3 $LOKI/tools/policy-load.py --as-args
358
+ --require-receipt
359
+ ```
360
+
361
+ With no policy file present it is explicit about what that means:
362
+
363
+ ```
364
+ $ python3 $LOKI/tools/policy-load.py
365
+ policy-load: no policy file at .loki-policy.json: a gate with no policy enforces nothing
366
+ ```
367
+ Exit 1. `--file` names a different path; `--json` emits the validated policy.
368
+
369
+ ### `policy-diff.py`
370
+
371
+ **Answers:** does this policy edit tighten the gate or weaken it? A loosening
372
+ looks like any other line in a diff, so this classifies by safety direction.
373
+
374
+ ```
375
+ $ python3 tools/policy-diff.py old-policy.json new-policy.json
376
+ WEAKENS: max_usd: 5.0 -> 20.0
377
+ 1 weakening(s), 0 unknown direction, 1 change(s) total
378
+ ```
379
+
380
+ `--fail-on-weaken` makes that a blocking review step:
381
+
382
+ ```
383
+ $ python3 tools/policy-diff.py old-policy.json new-policy.json --fail-on-weaken
384
+ WEAKENS: max_usd: 5.0 -> 20.0
385
+ policy-diff: 1 weakening(s) require an explicit human ack
386
+ ```
387
+ Exit 1.
388
+
389
+ ### `gate-report.py`
390
+
391
+ **Answers:** how do I show this verdict in CI-native form? Reads a `ci-gate --json`
392
+ verdict on stdin and re-renders it. It never invents a verdict.
393
+
394
+ ```text
395
+ $ cd /path/to/your/workspace
396
+ $ python3 $LOKI/tools/ci-gate.py . --require-receipt --json | python3 $LOKI/tools/gate-report.py --format markdown
397
+ ## Merge gate: PASS
398
+
399
+ | Policy | State | Detail |
400
+ | --- | --- | --- |
401
+ | receipt | PASS | VERIFIED |
402
+
403
+ PASS -- 1 of 1 policies passed
404
+ ```
405
+ (the output is markdown; it is indented here only so it does not render as a
406
+ heading on this page)
407
+
408
+ `--format github` emits workflow annotations instead. Shown here on a FAILING
409
+ verdict, since a passing gate has no annotations worth reading:
410
+
411
+ ```
412
+ ::error title=gate: receipt (FAIL)::FAILED
413
+ ::error title=merge gate FAIL::0 of 1 policies passed
414
+ ```
415
+ Formats: `markdown` (step summary), `github` (annotations), `text`. `--file`
416
+ attaches a path to github annotations, omitted when not given because an
417
+ annotation on the wrong file is worse than none.
418
+
419
+ ### `gate-explain.py`
420
+
421
+ **Answers:** the gate failed, so what do I run next? Turns a verdict into an
422
+ actionable command. On a FAILING verdict:
423
+
424
+ ```
425
+ $ python3 $LOKI/tools/ci-gate.py . --require-receipt --json | python3 $LOKI/tools/gate-explain.py
426
+ POLICY: receipt [FAIL]
427
+ CHECKED: the gate checked this and it failed
428
+ FOUND: FAILED
429
+ NEXT: attestation was required and did not hold. Read the attestation's own per-axis states before changing anything:
430
+ python3 tools/receipt-attest.py <workspace>/.loki/proofs/*/proof.json --json
431
+
432
+ GATE: FAIL -- 0 of 1 policies passed
433
+ ```
434
+
435
+ And on a passing one, it says so rather than manufacturing advice:
436
+
437
+ ```
438
+ POLICY: receipt [PASS]
439
+ CHECKED: the gate checked this and it passed
440
+ FOUND: VERIFIED
441
+ NEXT: nothing to do
442
+
443
+ GATE: PASS -- 1 of 1 policies passed
444
+ ```
445
+
446
+ ### `gate-badge.py`
447
+
448
+ **Answers:** what is the live gate state, as a README badge? Emits shields.io
449
+ endpoint JSON without laundering a failure into a green badge.
450
+
451
+ ```
452
+ $ python3 $LOKI/tools/ci-gate.py . --require-receipt --json | python3 $LOKI/tools/gate-badge.py
453
+ {"schemaVersion": 1, "label": "gate", "message": "passing", "color": "brightgreen"}
454
+ ```
455
+
456
+ On a failing verdict the same command emits, and the badge exit code follows the
457
+ verdict rather than always succeeding:
458
+
459
+ ```
460
+ {"schemaVersion": 1, "label": "gate", "message": "failing", "color": "red"}
461
+ ```
462
+ `--label` changes the badge label.
463
+
464
+ ### `gate-log.py`
465
+
466
+ **Answers:** one verdict is a fact; what is the pattern across a hundred? A gate
467
+ that was never recorded is not a gate that never blocked.
468
+
469
+ ```
470
+ $ cd /path/to/your/workspace
471
+ $ python3 $LOKI/tools/ci-gate.py . --require-receipt --json | python3 $LOKI/tools/gate-log.py record
472
+ gate-log: recorded pass to .loki/gate-log.jsonl
473
+
474
+ $ python3 $LOKI/tools/gate-log.py report
475
+ gate-log: 1 record(s)
476
+
477
+ pass 1
478
+ fail 0
479
+ unevaluable (NOT a pass) 0
480
+ corrupt 0
481
+
482
+ most failing policy: none -- no FAIL in 1 readable record(s)
483
+ trend: UNKNOWN -- fewer than 2 readable records (1)
484
+ ```
485
+ Note that `unevaluable` is counted separately and never folded into `pass`.
486
+
487
+ With no log at all it reports UNKNOWN, not a clean history:
488
+
489
+ ```
490
+ $ python3 $LOKI/tools/gate-log.py report
491
+ gate-log: no log at .loki/gate-log.jsonl -- UNKNOWN, not a clean history. A gate that was never recorded is not a gate that never blocked.
492
+ ```
493
+ Exit 66. The log path is relative to cwd, so `record` and `report` must run from
494
+ the same directory.
495
+
496
+ ---
497
+
498
+ ## Govern cost
499
+
500
+ Every tool here reports UNKNOWN when cost was not measured. None of them
501
+ substitutes zero for an absent measurement.
502
+
503
+ ### `preflight.sh`
504
+
505
+ **Answers:** will this run succeed, and what will it cost? Answered BEFORE the
506
+ run. Read-only.
507
+
508
+ ```
509
+ $ bash tools/preflight.sh /tmp/ws
510
+ Loki preflight -- /tmp/ws
511
+ Read-only: starts nothing, spends nothing, contacts no provider.
512
+
513
+ Provider
514
+ OK auto-detection would select: claude
515
+
516
+ Environment
517
+ OK doctor: 12 checks passed, 1 warning(s)
518
+
519
+ Saved state
520
+ no resumable run found
521
+
522
+ Projected cost
523
+ UNKNOWN no measured, priced iteration in this workspace
524
+ no iteration records found in this workspace: there is NO history to project from, so no cost is estimated (not $0.00)
525
+
526
+ ========================================
527
+ VERDICT: READY
528
+ ```
529
+ `--iterations N` projects over N iterations; without it the total reads UNKNOWN
530
+ rather than being guessed.
531
+
532
+ ### `cost-guard.py`
533
+
534
+ **Answers:** did this run's cost regress past a budget policy? A cost gate for CI.
535
+
536
+ ```
537
+ $ python3 tools/cost-guard.py /tmp/ws --max-usd 5
538
+ CANNOT EVALUATE: cost is UNMEASURED for /tmp/ws -- no efficiency record carried an observed cost or token count. Unmeasured is not within budget: this gate cannot say whether the run complied, so it reports no verdict rather than a green one.
539
+ ```
540
+ Exit 2. This is the convention's sharpest case: an unmeasured run is not a
541
+ passing run. `--baseline` plus `--max-increase-pct` gates on relative growth
542
+ instead of an absolute ceiling.
543
+
544
+ ### `token-guard.py`
545
+
546
+ **Answers:** did token usage regress? A provider-independent work ceiling, for
547
+ when you want a limit that does not move with pricing.
548
+
549
+ ```
550
+ $ python3 tools/token-guard.py /tmp/ws --max-output-tokens 100000
551
+ CANNOT EVALUATE: tokens are UNMEASURED for /tmp/ws -- no efficiency record carried an observed token count. Unmeasured is not within budget: this gate cannot say whether the run complied, so it reports no verdict rather than a green one.
552
+ ```
553
+ Exit 2. `--max-output-tokens` bounds output alone (the work signal);
554
+ `--max-total-tokens` bounds input + output + cache reads, which cache reads
555
+ dominate.
556
+
557
+ ### `baseline-pin.py`
558
+
559
+ **Answers:** which run is THE cost baseline? Pin one, resolve it later for
560
+ `cost-guard --baseline`.
561
+
562
+ ```
563
+ $ cd /path/to/your/workspace
564
+ $ python3 $LOKI/tools/baseline-pin.py set .
565
+ REFUSED TO PIN: receipt ./.loki/proofs/good/proof.json records NO MEASURED COST, so it must not become a baseline: a percentage increase against an unmeasured number is undefined. Unmeasured is not $0.00. Fix the cost instrumentation for that run, then pin it.
566
+ ```
567
+ Exit 1. Subcommands: `set` pins the newest receipt in a workspace, `show` reports
568
+ what is pinned and whether it is intact, `path` prints the pinned proof path for
569
+ feeding to `--baseline`.
570
+
571
+ ### `cost-history.py`
572
+
573
+ **Answers:** what is the cost trend across MANY runs? Unmeasured runs are stored
574
+ as null and excluded from the trend, never counted as zero.
575
+
576
+ ```
577
+ $ cd /path/to/your/workspace
578
+ $ python3 $LOKI/tools/cost-history.py record .
579
+ recorded /private/var/folders/.../loki-doccheck-cayf1yk6: UNMEASURED (stored as null, excluded from the trend)
580
+
581
+ $ python3 $LOKI/tools/cost-history.py report
582
+ 1 run(s), 0 measured
583
+ NO TREND: 1 run(s) recorded but none carried a measured cost; there is nothing to trend. Unmeasured runs are kept as null, not 0.
584
+ ```
585
+
586
+ ### `estimate-run.py`
587
+
588
+ **Answers:** what is this run likely to cost, and on what basis? Projects from
589
+ measured history only.
590
+
591
+ ```
592
+ $ python3 tools/estimate-run.py /tmp/ws --iterations 5
593
+ Run cost ESTIMATE -- /tmp/ws
594
+
595
+ NO BASIS: no measured, priced iteration to project from.
596
+ Cost per iteration: UNKNOWN
597
+ Projected cost: UNKNOWN
598
+ Records found: 0 measured: 0 priced: 0
599
+ Basis model(s): not recorded
600
+ Model now: not pinned
601
+
602
+ - no iteration records found in this workspace: there is NO history to project from, so no cost is estimated (not $0.00)
603
+ ```
604
+ Exit 0: "no basis" is an honest answer to the question asked, and no gate
605
+ consumes this exit code.
606
+
607
+ ### `cost-attribute.py`
608
+
609
+ **Answers:** where did a run's cost and time actually GO, stage by stage?
610
+
611
+ ```
612
+ $ python3 tools/cost-attribute.py /tmp/ws
613
+ NO DATA: no events at /tmp/ws/.loki/events.jsonl -- pass the workspace directory of a completed run
614
+ ```
615
+ Exit 66. Needs `.loki/events.jsonl` from a completed run.
616
+
617
+ ### `model-advisor.py`
618
+
619
+ **Answers:** would a cheaper model have done this job, and what would it have
620
+ saved? An advisor, deliberately not a gate.
621
+
622
+ ```
623
+ $ python3 tools/model-advisor.py /tmp/ws
624
+ Model cost advisor -- /tmp/ws
625
+
626
+ NO BASIS: no measured, priced iteration in this workspace.
627
+ Model used: UNKNOWN
628
+ Measured cost: UNKNOWN
629
+ Recommendation: NONE -- there is no measured basis
630
+ Projected saving: UNKNOWN
631
+ Records found: 0 measured: 0 priced: 0
632
+
633
+ Cited external benchmark (SWE-bench verified) -- NOT a measurement of your workload:
634
+ MiniMax M2.5 (open weights) score 75.8 cost $36.64
635
+ Claude Opus 4.6 score 75.6 cost $275.76
636
+ the harness itself was worth about 3.4 points on an identical model
637
+ equal-or-better score at roughly 7.5x lower cost ON THAT BENCHMARK. It is not
638
+ a measurement of your workload and does not predict that a cheaper model would
639
+ complete YOUR task
640
+ ```
641
+ Exit 0 with no basis, by design. `tests/test_tool_exit_contract.py` pins this
642
+ judgement so a future reader does not "fix" it: forcing an advisor non-zero
643
+ trains operators to ignore a failing advisor.
644
+
645
+ ---
646
+
647
+ ## Diagnose a run
648
+
649
+ ### `run-replay.py`
650
+
651
+ **Answers:** what did a completed run actually do, iteration by iteration?
652
+ Reconstructs from artifacts; starts nothing, spends nothing.
653
+
654
+ ```
655
+ $ python3 tools/run-replay.py /tmp/ws
656
+ no events at /tmp/ws/.loki/events.jsonl -- pass the workspace directory of a completed run
657
+ ```
658
+ Exit 66. Point it at the workspace of a run that finished.
659
+
660
+ ### `tool-index.py`
661
+
662
+ **Answers:** what tools shipped, and which can a user reach?
663
+
664
+ ```
665
+ $ python3 tools/tool-index.py
666
+ loki tools
667
+
668
+ baseline-pin.py Pin one run as THE cost baseline, then resolve it later.
669
+ ci-gate.py One exit code for EVERY merge policy. The gate a CI job actually calls.
670
+ cost-attribute.py Where did a run's cost and time actually GO, stage by stage.
671
+ ...
672
+ ```
673
+ `--json` adds a `shipped` flag per tool. `--tools-dir` points at a different
674
+ directory.
675
+
676
+ ### `hybrid_search.py`
677
+
678
+ **Answers:** where in this codebase is X? Merges lexical (grep/ripgrep) and
679
+ semantic search under a token budget.
680
+
681
+ ```
682
+ $ python3 tools/hybrid_search.py council_should_stop --grep-only --top 3
683
+ hybrid search: 'council_should_stop' [grep-only] (ripgrep not found, using grep)
684
+ budget: 3000 tokens, 3 result(s)
685
+
686
+ [1] tests/test-council-convergence-floor.sh:223 (match: grep, score: 0.016393)
687
+ # The HARD floor in council_should_stop (ITERATION_COUNT < floor -> not allowed to
688
+ ```
689
+ `--semantic-only` requires the codebase index (see `index-codebase.py` below,
690
+ which currently does not run on Python 3.14). `--grep-only` needs nothing.
691
+
692
+ ---
693
+
694
+ ## Internal and maintenance
695
+
696
+ Five of the tools below are excluded from the exit-code convention **by name** in
697
+ `tests/test_tool_exit_contract.py` (the sixth named exclusion, `hybrid_search.py`,
698
+ is user-facing and appears under [Diagnose a run](#diagnose-a-run)). They are
699
+ benchmark harnesses and repo maintenance scripts, not CI-composable gates.
700
+
701
+ `verify-demo.sh` is listed here for a different reason: it is user-facing and is
702
+ documented in full under [Verify a receipt](#verify-a-receipt), but as a shell
703
+ script it is never globbed by the contract test, so nothing enforces its exit
704
+ codes.
705
+
706
+ Several of these carry version-stamped internal release notes as their
707
+ docstrings rather than user-facing descriptions, so what follows is a description
708
+ of observed behaviour, not a quoted one.
709
+
710
+ ### `bench_memory_retrieval.py`
711
+
712
+ Memory retrieval cold-start benchmark. Seeds N episodes, performs cold
713
+ retrievals, reports percentiles against a threshold.
714
+
715
+ ```
716
+ $ python3 tools/bench_memory_retrieval.py --episodes 50 --runs 5 --json
717
+ {
718
+ "episodes_seeded": 50,
719
+ "runs": 5,
720
+ "seed_ms": 47.5,
721
+ "p50_ms": 4.7,
722
+ "p95_ms": 5.2,
723
+ "p99_ms": 5.2,
724
+ "threshold_ms": 500.0,
725
+ "p95_under_threshold": true,
726
+ "generated_at": "2026-08-03T02:38:51.517506+00:00"
727
+ }
728
+ ```
729
+ Defaults are 1000 episodes, 100 runs, 500ms threshold. Its own `--help` notes
730
+ that 10000 episodes does NOT meet the 500ms bar with file-based storage.
731
+
732
+ ### `bench_cross_project_lift.py`
733
+
734
+ Measures how much retrieval coverage a project gains from sibling projects'
735
+ memory. The `method` field names its own limitation.
736
+
737
+ ```
738
+ $ python3 tools/bench_cross_project_lift.py --json
739
+ {
740
+ "goals": 6,
741
+ "baseline_covered": 0,
742
+ "cross_covered": 3,
743
+ "lift_absolute": 3,
744
+ "lift_pct_points": 50.0,
745
+ "net_new_from_siblings": 3,
746
+ "top_k": 5,
747
+ "method": "retrieval-coverage (keyword-overlap relevance proxy), NOT task-success"
748
+ }
749
+ ```
750
+
751
+ ### `probe-model-catalog.py`
752
+
753
+ Fetches provider documentation pages, extracts model IDs by regex, and compares
754
+ them against `providers/model_catalog.json`. It reports and never auto-rewrites
755
+ the catalog.
756
+
757
+ ```
758
+ $ python3 tools/probe-model-catalog.py
759
+ == claude ==
760
+ known in catalog: 4
761
+ found in docs: 14
762
+ NEW CANDIDATES: claude-haiku-4-5-20251001, claude-opus-4-1, ...
763
+
764
+ To adopt a new model: edit providers/model_catalog.json -> bump latest_<tier>
765
+ and add to models[]. Then re-run this script to confirm it disappears from new_candidates.
766
+ ```
767
+ `--strict` exits nonzero when new models are found. Note it still probes a
768
+ `gemini` section; Gemini was removed as a provider, so those candidates are not
769
+ adoptable.
770
+
771
+ ### `regen-state-machine-refs.py`
772
+
773
+ Verifies that every `file:line (function)` reference in
774
+ `docs/architecture/STATE-MACHINES.md` still points at that function's definition.
775
+ `--fix` rewrites stale numbers in place, `--strict` exits nonzero on drift (for
776
+ CI).
777
+
778
+ ### `index-codebase.py`
779
+
780
+ Builds the semantic index that `hybrid_search.py --semantic-only` consumes.
781
+
782
+ **This does not currently run on this machine.** It imports `chromadb`, which
783
+ fails on Python 3.14:
784
+
785
+ ```
786
+ $ python3 tools/index-codebase.py --help
787
+ Traceback (most recent call last):
788
+ File ".../tools/index-codebase.py", line 41, in <module>
789
+ import chromadb
790
+ ...
791
+ UserWarning: Core Pydantic V1 functionality isn't compatible with Python 3.14 or greater.
792
+ ```
793
+ Exit 1 even for `--help`. Use `hybrid_search.py --grep-only` until the
794
+ dependency supports your interpreter. No usage example is given here because
795
+ none could be executed.
796
+
797
+ ### `verify-demo.sh`
798
+
799
+ Documented under [Verify a receipt](#verify-a-receipt) above; it is user-facing
800
+ despite also serving as a smoke test.