loki-mode 9.8.0 → 9.12.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +19 -14
- package/SKILL.md +3 -2
- package/VERSION +1 -1
- package/autonomy/loki +122 -1
- package/autonomy/run.sh +49 -2
- package/dashboard/__init__.py +1 -1
- package/dashboard/api_evidence.py +411 -0
- package/dashboard/api_operator.py +283 -0
- package/dashboard/api_phases.py +262 -0
- package/dashboard/api_releases.py +242 -0
- package/dashboard/api_runs.py +477 -0
- package/dashboard/api_tests.py +444 -0
- package/dashboard/api_v2.py +47 -1
- package/dashboard/server.py +54 -0
- package/dashboard/static/index.html +246 -135
- package/docs/ARCHITECTURE-OVERVIEW.md +5 -3
- package/docs/CAPABILITY-BACKLOG.md +53 -0
- package/docs/COMPARISON.md +2 -2
- package/docs/COMPETITIVE-ANALYSIS.md +1 -1
- package/docs/COMPETITIVE-SCORECARD.md +422 -0
- package/docs/DASHBOARD-9.12-EVIDENCE.md +97 -0
- package/docs/DASHBOARD-ARCHITECTURE.md +423 -0
- package/docs/DEMOS.md +21 -23
- package/docs/HANDOFF-2026-08-03.md +439 -0
- package/docs/INSTALLATION.md +17 -10
- package/docs/OUTCOME-FRONTIER.md +536 -0
- package/docs/PROMPT-ABLATION-RESULT.md +97 -0
- package/docs/TOOLS.md +800 -0
- package/docs/alternative-installations.md +2 -3
- package/docs/audit-logging.md +44 -35
- package/docs/authentication.md +13 -2
- package/docs/authorization.md +87 -81
- package/docs/git-workflow.md +6 -3
- package/docs/metrics.md +15 -16
- package/docs/network-security.md +16 -13
- package/docs/openclaw-integration.md +36 -556
- package/docs/show-hn-post.md +2 -2
- package/docs/siem-integration.md +39 -36
- package/loki-ts/dist/loki.js +18 -18
- package/mcp/__init__.py +1 -1
- package/package.json +2 -2
- package/plugins/loki-mode/.claude-plugin/plugin.json +1 -1
- package/references/confidence-routing.md +18 -1
- package/references/invariant-checks.md +13 -8
- package/references/magic-rarv-integration.md +0 -1
- package/references/multi-provider.md +27 -5
- package/skills/healing.md +4 -2
- package/tools/audit-docs.py +488 -0
- package/tools/baseline-pin.py +19 -1
- package/tools/calibration-audit.py +523 -0
- package/tools/ci-gate.py +19 -1
- package/tools/cost-forecast.py +344 -0
- package/tools/cost-guard.py +19 -1
- package/tools/cost-history.py +19 -1
- package/tools/cost-per-outcome.py +394 -0
- package/tools/estimate-run.py +19 -1
- package/tools/evidence-freshness.py +307 -0
- package/tools/gate-init.py +19 -1
- package/tools/gate-report.py +19 -1
- package/tools/gate-simulate.py +570 -0
- package/tools/gate-trend.py +354 -0
- package/tools/model-advisor.py +52 -1
- package/tools/policy-load.py +19 -1
- package/tools/prompt-cost.py +363 -0
- package/tools/prompt-diff.py +448 -0
- package/tools/prompt-lint.py +448 -0
- package/tools/receipt-bundle.py +72 -2
- package/tools/receipt-diff.py +19 -1
- package/tools/receipt-find.py +19 -1
- package/tools/receipt-stats.py +380 -0
- package/tools/receipt-timeline.py +478 -0
- package/tools/receipt-verify-batch.py +291 -0
- package/tools/run-replay.py +19 -1
- package/tools/signing-status.py +19 -1
- package/tools/token-guard.py +19 -1
- package/tools/token-tax.py +375 -0
- package/tools/tool-index.py +19 -1
- package/tools/verification-tax.py +277 -0
- package/tools/verify-chain.py +361 -0
|
@@ -0,0 +1,439 @@
|
|
|
1
|
+
# Handoff, 2026-08-03
|
|
2
|
+
|
|
3
|
+
## 0. DECISION-GRADE RC SUMMARY (read this first)
|
|
4
|
+
|
|
5
|
+
Pinned: **`worktree-pre-push-scoped-pytest` @ `16ba5c5e`**.
|
|
6
|
+
|
|
7
|
+
| Fact | Value | Command |
|
|
8
|
+
|---|---|---|
|
|
9
|
+
| RC HEAD | `16ba5c5e` at the time of writing; this document's own commit
|
|
10
|
+
advances it, so ALWAYS trust the command over this cell | `git rev-parse HEAD` |
|
|
11
|
+
| vs `origin/main` (`dd692561`) | **27 ahead, 0 behind** | `git rev-list --count origin/main..HEAD` |
|
|
12
|
+
| vs RC remote (`3808f63f`) | **17 ahead, 0 behind** | `git rev-list --count origin/worktree-...` |
|
|
13
|
+
|
|
14
|
+
WHY THE SHA CELL IS ALWAYS ONE BEHIND. Committing this file moves HEAD, so
|
|
15
|
+
a handoff can never pin its own commit. The cell records the SHA the
|
|
16
|
+
receipts below were taken on; the live value is one commit later. Re-run
|
|
17
|
+
`git rev-parse --short HEAD` and `git rev-list --count origin/main..HEAD`
|
|
18
|
+
rather than trusting any number typed here -- that habit is what caught
|
|
19
|
+
three stale pins in this document already.
|
|
20
|
+
|
|
21
|
+
LOCAL EVIDENCE AND CI EVIDENCE ARE DIFFERENT THINGS, and this document keeps
|
|
22
|
+
them apart deliberately. Everything under "Local receipts" ran on this machine
|
|
23
|
+
against this exact SHA. NO CI has ever run on `16ba5c5e` or on any commit in this
|
|
24
|
+
branch: a branch push triggers no workflow (see gate 2), so there is no CI
|
|
25
|
+
verdict to report for the RC. The CI table below is for `origin/main` only.
|
|
26
|
+
|
|
27
|
+
**A BRANCH PUSH ALONE RUNS NO CI.** Every workflow triggers only on
|
|
28
|
+
push-to-main or a pull_request targeting main, and `test.yml` has no
|
|
29
|
+
`workflow_dispatch`. Pushing this branch produces zero runs, as observed.
|
|
30
|
+
A CI verdict on this SHA requires a branch push AND a PR to main.
|
|
31
|
+
|
|
32
|
+
### CI evidence, exact
|
|
33
|
+
|
|
34
|
+
`origin/main` = `dd692561`:
|
|
35
|
+
|
|
36
|
+
| Workflow | Conclusion | URL |
|
|
37
|
+
|---|---|---|
|
|
38
|
+
| Post-Release Soak Monitor | success | actions/runs/30805237104 |
|
|
39
|
+
| Security Audit | success | actions/runs/30796434331 |
|
|
40
|
+
| **Tests** | **failure** | actions/runs/30794017151 |
|
|
41
|
+
|
|
42
|
+
The Tests failure is ONE job, ONE cause: `Shell tests (shard 3/4)` ->
|
|
43
|
+
`ShellCheck Linting FAILED`. Shards 0, 1, 2 pass. Held commit `a4d8b839`
|
|
44
|
+
takes shellcheck from 7 warning sites to **361 passed / 0 failed**.
|
|
45
|
+
|
|
46
|
+
**RC branch CI: 0 runs, and it will stay 0.** Every workflow triggers only on
|
|
47
|
+
push-to-main or a PR targeting main, and `test.yml` has no `workflow_dispatch`.
|
|
48
|
+
A branch push cannot produce a CI verdict. PR #181 is MERGED 2026-07-28 on a
|
|
49
|
+
different diff and is NOT evidence about current main.
|
|
50
|
+
|
|
51
|
+
### Local receipts on the exact RC HEAD `16ba5c5e`
|
|
52
|
+
|
|
53
|
+
Measured on this machine. NOT CI: no workflow has run on this SHA.
|
|
54
|
+
|
|
55
|
+
| Check | Result |
|
|
56
|
+
|---|---|
|
|
57
|
+
| `bash scripts/local-ci.sh` | **exit 0**, Failed 0 |
|
|
58
|
+
| `bash tests/run-shellcheck.sh` | **361 passed, 0 failed** |
|
|
59
|
+
| `bash tests/test-plan-command.sh` | **27 passed, 0 failed** (was 25/2; two stale June cost pins fixed) |
|
|
60
|
+
| `python3 -m pytest tests/ -q -k "bench or schema"` | **149 passed**, 1 skipped |
|
|
61
|
+
| `python3 -m pytest tests/` (full) | **2970 passed, 1 failed**, 19 skipped |
|
|
62
|
+
|
|
63
|
+
THE ONE FULL-SUITE FAILURE, characterised rather than hidden:
|
|
64
|
+
`tests/dashboard/test_build_supervisor.py::test_confined_claude_auth_uses_exact_login_keychain_capability`.
|
|
65
|
+
|
|
66
|
+
- PRE-EXISTING, not caused by anything here: a pristine `git archive HEAD`
|
|
67
|
+
extract with no working-tree residue reproduces it.
|
|
68
|
+
- ENVIRONMENT-SENSITIVE, not a code defect: the same file passes **37/37 in
|
|
69
|
+
isolation in this worktree, twice in a row**, and the failing assertion is
|
|
70
|
+
about a real login-keychain capability, which differs between this worktree
|
|
71
|
+
and a bare extract.
|
|
72
|
+
- It is NOT in the fast tier, which is why `local-ci.sh` is green while the
|
|
73
|
+
full suite is not. Both numbers are reported rather than the flattering one.
|
|
74
|
+
|
|
75
|
+
### Working-tree residue, classified. NOTHING deleted, reverted or committed.
|
|
76
|
+
|
|
77
|
+
| Path | Class | Note |
|
|
78
|
+
|---|---|---|
|
|
79
|
+
| `coverage/clover.xml` | generated residue | rewritten by test runs |
|
|
80
|
+
| `coverage/lcov-report/index.html` | generated residue | rewritten by test runs |
|
|
81
|
+
| `loki-ts/dist/loki.js.map` | generated residue | rewritten by `bun run build` |
|
|
82
|
+
| `f.txt` (deleted) | **user-owned** | deleted in working tree only; **still present in HEAD**, fully recoverable via `git checkout -- f.txt` |
|
|
83
|
+
| `benchmarks/results/prompt-ablation.jsonl` | measurement output | untracked; the 6 real ablation trials |
|
|
84
|
+
|
|
85
|
+
|
|
86
|
+
### Every held commit (27, unpushed to main)
|
|
87
|
+
|
|
88
|
+
| SHA | Subject |
|
|
89
|
+
|---|---|
|
|
90
|
+
| `16ba5c5e` | feat(bench): `loki bench oracles` -- reachable where the spend decision is made |
|
|
91
|
+
| `5cc5ffd2` | feat(bench): attest the artifact hashes so an edited answer key cannot pass |
|
|
92
|
+
| `6dbef513` | docs: both remaining tiers now verified against a near-miss, gate 3 precondition MET |
|
|
93
|
+
| `89babec7` | docs: gate 3's precondition is PARTIALLY met, with the hashes that make it checkable |
|
|
94
|
+
| `43e4d530` | test(bench): pin the private probe's own guarantees, and refuse to promote a tier on private evidence |
|
|
95
|
+
| `f9895095` | feat(bench): hash-bound private oracle probes -- the replayable version of what I retracted |
|
|
96
|
+
| `ed9f7a6c` | fix(bench,docs): retract the medium/high VERIFIED label -- the evidence cannot be replayed |
|
|
97
|
+
| `a54c190c` | feat(bench): medium and high oracles verified against a plausible wrong answer |
|
|
98
|
+
| `1fd65420` | docs: make HEAD decision-ready -- exact-SHA receipts, local evidence kept apart from CI |
|
|
99
|
+
| `5e0ed5b5` | fix(tests): two demo assertions pinned a June cost the estimator has since corrected |
|
|
100
|
+
| `49ebefa2` | docs: reconcile the handoff to live topology and separate the three founder gates |
|
|
101
|
+
| `7b25a772` | docs(bench): record that adversarially_unverified is unreachable today, and why it stays |
|
|
102
|
+
| `ea930bbb` | fix(bench,ci): fail on model mismatch, wire the integrity check, add adversarial oracle probes |
|
|
103
|
+
| `8980313e` | docs: decision-grade RC handoff reconciled to current topology |
|
|
104
|
+
| `2c031a14` | feat(safety): default-deny release interlock, and a detector for the recurring core.bare fault |
|
|
105
|
+
| `6ffa5080` | feat(bench): founder-gated spend interlock, and keep grader evidence beside the measured flag |
|
|
106
|
+
| `35add6d7` | fix(bench): flag acceptance commands that cannot fail -- a timeout scored as success |
|
|
107
|
+
| `3808f63f` | feat(bench): cross-tier oracle receipt -- verifies the graders, runs no agent |
|
|
108
|
+
| `1591f880` | docs: re-verify ref topology after an audit, and name the single CI blocker |
|
|
109
|
+
| `ef0a6d49` | docs: put the spend decision's actual numbers in the handoff |
|
|
110
|
+
| `12669cd3` | docs: the no-spend outcome frontier -- 3 held-out tasks, deterministic oracles, nothing run |
|
|
111
|
+
| `fdba7023` | docs: handoff -- reconciliation, receipts, risks, founder-gated next action |
|
|
112
|
+
| `19c2b5ad` | fix(tests): bound each completion probe, so a hang is a NAMED failure not rc143 |
|
|
113
|
+
| `90468210` | feat(advisor): point at the calibration signal, and pin that it can never rank |
|
|
114
|
+
| `0f4da110` | docs: freeze the T1/T2 benchmark designs, unrun and unpaid |
|
|
115
|
+
| `8c33d123` | feat(tools): offline calibration audit, and it discloses that it cannot measure accuracy |
|
|
116
|
+
| `a4d8b839` | fix(lint): the last red CI job was three pre-existing shellcheck warnings |
|
|
117
|
+
|
|
118
|
+
### THE THREE FOUNDER GATES, deliberately separate
|
|
119
|
+
|
|
120
|
+
Ordered by dependency, not by value. Gate 1 first because gates 2 and 3 both
|
|
121
|
+
move code toward main, and today nothing non-bypassable stands between main
|
|
122
|
+
and a publish.
|
|
123
|
+
|
|
124
|
+
**GATE 1 -- a durable, non-bypassable publish control.** Zero spend, zero
|
|
125
|
+
code. Either an `environment:` key on release.yml's publish jobs (GitHub then
|
|
126
|
+
requires a named reviewer before those jobs start), or a ruleset on main.
|
|
127
|
+
Needs repository settings access, which is why it is not done here.
|
|
128
|
+
|
|
129
|
+
Everything I built is operator-side and BYPASSABLE by not running it:
|
|
130
|
+
`scripts/release-approval-gate.sh` refuses by default when run, and is a
|
|
131
|
+
no-op when skipped. Treat a green run as "the operator checked", never as
|
|
132
|
+
"publishing is gated". Measured today: release.yml triggers on any
|
|
133
|
+
VERSION-path push to main and runs `gh release create`, `npm publish` twice
|
|
134
|
+
and a Docker push, with no `environment:` key, and
|
|
135
|
+
`gh api .../branches/main/protection` returns 404.
|
|
136
|
+
|
|
137
|
+
**GATE 2 -- branch push PLUS a PR to main, for CI on this exact SHA.** Zero
|
|
138
|
+
spend. These are ONE gate because either alone yields nothing: a branch push
|
|
139
|
+
runs no workflow, and there is no PR without the push. Only this produces a
|
|
140
|
+
CI verdict on `43e4d530`. It is also the only path by which held `a4d8b839`
|
|
141
|
+
reaches main and clears the single remaining Tests failure.
|
|
142
|
+
|
|
143
|
+
**GATE 3 -- benchmark spend, roughly $25-$180 and 4-12 hours.** Independent
|
|
144
|
+
of gates 1 and 2. Buys the first OUTCOME evidence rather than affordance
|
|
145
|
+
evidence (`docs/OUTCOME-FRONTIER.md`).
|
|
146
|
+
|
|
147
|
+
PRECONDITION ON GATE 3: **MET, with one qualifier that decides who can
|
|
148
|
+
check it.** Small is verified by a receipt ANY reader can reproduce.
|
|
149
|
+
Medium and high are verified by a PRIVATE hash-bound receipt (below):
|
|
150
|
+
replayable by whoever holds the artifacts, not by a reader on a fresh
|
|
151
|
+
clone.
|
|
152
|
+
|
|
153
|
+
| Tier | Positive | Adversarial | Receipt status |
|
|
154
|
+
|---|---|---|---|
|
|
155
|
+
| small `simple-1-contact-form` | exit 0 | exit 1 | **VERIFIED, reproducible** -- fixtures are committed, `tier_receipt.py` re-runs both probes on demand |
|
|
156
|
+
| medium `multifail-1-two-modules` | -- | -- | **NOT ATTEMPTED** in the shared receipt; VERIFIED privately, see below |
|
|
157
|
+
| high `hard-1-order-api` | -- | -- | **NOT ATTEMPTED** in the shared receipt; VERIFIED privately, see below |
|
|
158
|
+
|
|
159
|
+
NON-REPRODUCIBLE OBSERVATION, recorded as a NOTE and explicitly NOT a
|
|
160
|
+
receipt. On 2026-08-03 the medium and high graders were exercised by hand
|
|
161
|
+
against artifacts built in a temp directory and destroyed. Observed exits
|
|
162
|
+
were 0 for a correct implementation and 1 for a near-miss (high: omitting
|
|
163
|
+
the subtotal>=100 discount; medium: fixing only cluster A). The high-tier
|
|
164
|
+
near-miss passed all ten validation cases and both non-discount totals.
|
|
165
|
+
|
|
166
|
+
That observation must not be promoted to VERIFIED, and the earlier draft of
|
|
167
|
+
this section wrongly did. The artifacts are gone, no hash binds them to the
|
|
168
|
+
run, and nothing in this repo can replay it -- so it fails the same standard
|
|
169
|
+
this codebase applies to every other claim: an unreproducible observation is
|
|
170
|
+
not evidence. It is recorded because it is a useful prior for whoever builds
|
|
171
|
+
the real probes, and for no stronger purpose.
|
|
172
|
+
|
|
173
|
+
THE MECHANISM THAT MEETS THIS PRECONDITION NOW EXISTS, and one of the two
|
|
174
|
+
tiers is verified against it. `benchmarks/bench/private_probe.py` runs the
|
|
175
|
+
in-repo graders against artifacts held OUTSIDE this repository (default
|
|
176
|
+
`~/loki-bench-private`), content-hashes each set over its sorted
|
|
177
|
+
(filename, bytes) pairs, and records the hash beside the exit code. The
|
|
178
|
+
artifacts stay out of the repo, so the held-out design is uncontaminated;
|
|
179
|
+
the hash makes the claim replayable, which the earlier by-hand run was not.
|
|
180
|
+
|
|
181
|
+
Measured 2026-08-03:
|
|
182
|
+
|
|
183
|
+
| Task | Probe | Exit | Expected | Artifact sha256 (first 16) |
|
|
184
|
+
|---|---|---|---|---|
|
|
185
|
+
| `hard-1-order-api` | positive | 0 | 0 | `884d555b6e53833d` |
|
|
186
|
+
| `hard-1-order-api` | adversarial | 1 | 1 | `486a62137c64167a` |
|
|
187
|
+
| `multifail-1-two-modules` | positive | 0 | 0 | `a7134c9c1ae36ae9` |
|
|
188
|
+
| `multifail-1-two-modules` | adversarial | 1 | 1 | `f4daff9d346c6116` |
|
|
189
|
+
|
|
190
|
+
Both tiers read `verified`. Each adversarial artifact is the near-miss the
|
|
191
|
+
grader's own docstring names: the high one omits the subtotal>=100 discount
|
|
192
|
+
while passing all ten validation cases and both non-discount totals; the
|
|
193
|
+
medium one fixes cluster A and leaves cluster B absent, and the grader
|
|
194
|
+
rejects it naming `cluster B_roman`.
|
|
195
|
+
|
|
196
|
+
So every tier now rejects a PLAUSIBLE WRONG answer, not merely an absent
|
|
197
|
+
one, which is the property that makes a paid run's number mean something.
|
|
198
|
+
|
|
199
|
+
HOW TO CHECK IT YOURSELF, and what stops it rotting.
|
|
200
|
+
|
|
201
|
+
loki bench oracles # or: python3 benchmarks/bench/private_probe.py
|
|
202
|
+
|
|
203
|
+
Reachable from the CLI deliberately, and placed beside `bench run`: the
|
|
204
|
+
question in front of a paid run is whether a green cell would mean
|
|
205
|
+
anything, so the answer sits one command from the spend decision. Costs
|
|
206
|
+
nothing -- local graders against artifacts already on disk, never a
|
|
207
|
+
provider.
|
|
208
|
+
|
|
209
|
+
`benchmarks/bench/private_attestation.json` commits the HASHES (never the
|
|
210
|
+
artifacts) and the probe compares against them. Without that the probe
|
|
211
|
+
would verify whatever bytes it found, so an artifact edited afterwards
|
|
212
|
+
would be re-blessed under a new hash while the receipt still read
|
|
213
|
+
`verified` -- which is how an edited answer key launders itself green.
|
|
214
|
+
|
|
215
|
+
Proven on the case that matters: appending a COSMETIC comment leaves
|
|
216
|
+
behaviour identical, the grader still returns its expected exit code, and
|
|
217
|
+
the run still fails with `attestation=DRIFTED` naming both hashes. Only
|
|
218
|
+
the hash comparison can catch that.
|
|
219
|
+
|
|
220
|
+
An unattested probe reads `unattested`, not drift. Absence of a record is
|
|
221
|
+
not evidence of tampering.
|
|
222
|
+
|
|
223
|
+
TWO QUALIFIERS THAT MUST TRAVEL WITH THAT TABLE.
|
|
224
|
+
|
|
225
|
+
First, this is a PRIVATE receipt. It is checkable by whoever holds the
|
|
226
|
+
artifacts, and not by a reader on a fresh clone. `tier_receipt.py`
|
|
227
|
+
therefore still reports medium and high as `not_attempted`, deliberately:
|
|
228
|
+
promoting them there would make a shared receipt's verdict depend on who
|
|
229
|
+
ran it, which re-creates the retracted defect one level up. The two
|
|
230
|
+
receipts answer different questions and a human combines them.
|
|
231
|
+
|
|
232
|
+
Second, the hashes above identify one operator's artifacts on one machine.
|
|
233
|
+
They prove the claim is REPLAYABLE, not that anyone else can replay it
|
|
234
|
+
today. Making it shareable (a second repository, an attested bundle) is a
|
|
235
|
+
distribution decision, not a code change.
|
|
236
|
+
|
|
237
|
+
### RELEASE EXPOSURE, measured today
|
|
238
|
+
|
|
239
|
+
`.github/workflows/release.yml` triggers on `push: paths: ['VERSION'],
|
|
240
|
+
branches: [main]` and runs `gh release create`, `npm publish --access public`
|
|
241
|
+
(twice) and a Docker push. It has **no `environment:` approval gate**, and
|
|
242
|
+
`gh api .../branches/main/protection` returns **404 Branch not protected**.
|
|
243
|
+
|
|
244
|
+
So a single VERSION-editing commit reaching main publishes with no human in
|
|
245
|
+
the loop. `scripts/release-approval-gate.sh` is a default-DENY local interlock
|
|
246
|
+
against exactly that, verified on a throwaway 9.11.0 -> 9.99.99 bump. The
|
|
247
|
+
durable fix (an `environment:` gate or a ruleset) changes the publishing path
|
|
248
|
+
or repo settings and is **founder-gated** -- it is listed as an approval still
|
|
249
|
+
needed, not silently applied.
|
|
250
|
+
|
|
251
|
+
|
|
252
|
+
State at handoff. Everything below is verified by a command whose output was
|
|
253
|
+
read, or is marked UNKNOWN. Nothing has been pushed, published, merged, or
|
|
254
|
+
released, and nothing was spent.
|
|
255
|
+
|
|
256
|
+
## 1. Reconciliation
|
|
257
|
+
|
|
258
|
+
SUPERSEDED SNAPSHOT. The HEAD and commit count below are from an earlier
|
|
259
|
+
point in the session and are kept as history, NOT as current state. Section 0
|
|
260
|
+
is the live topology: `7b25a772`, 16 ahead of origin/main, 6 ahead of the RC
|
|
261
|
+
remote. The VERSION / npm / tag rows are still accurate.
|
|
262
|
+
|
|
263
|
+
| Fact | Value (as of `19c2b5ad`) | How verified |
|
|
264
|
+
|---|---|---|
|
|
265
|
+
| Local HEAD | `19c2b5ad` (SUPERSEDED, now `7b25a772`) | `git rev-parse HEAD` |
|
|
266
|
+
| `origin/main` | `dd692561` (unchanged) | `git rev-parse origin/main` after fetch |
|
|
267
|
+
| Commits ahead | 5 (SUPERSEDED, now 16) | `git rev-list --count origin/main..HEAD` |
|
|
268
|
+
| repo VERSION | 9.11.0 | `cat VERSION` |
|
|
269
|
+
| npm latest | 9.8.1 | `npm view loki-mode version` |
|
|
270
|
+
| newest git tag | v9.8.1 | `git tag --sort=-v:refname \| head -1` |
|
|
271
|
+
|
|
272
|
+
**npm and the tag agree at 9.8.1.** The repo being at 9.11.0 is not drift: it
|
|
273
|
+
is an unreleased repo under a no-publish directive. Three minor versions of
|
|
274
|
+
work are committed and unpublished by instruction. Reconciled, not a defect.
|
|
275
|
+
|
|
276
|
+
## 1b. Ref topology, re-verified after an independent audit (09:04Z)
|
|
277
|
+
|
|
278
|
+
An audit reported the session ref `loki/session-1785167314-9337` at
|
|
279
|
+
`09138e26` as 0 ahead / 176 behind, fully contained in `origin/main`, and
|
|
280
|
+
concluded the "commits held/unpushed" claim was stale. **Both statements are
|
|
281
|
+
correct, about DIFFERENT refs.** Re-measured here:
|
|
282
|
+
|
|
283
|
+
```
|
|
284
|
+
git rev-list --left-right --count origin/main...loki/session-1785167314-9337
|
|
285
|
+
-> 176 0 (behind 176, ahead 0)
|
|
286
|
+
git merge-base --is-ancestor loki/session-1785167314-9337 origin/main
|
|
287
|
+
-> true origin/main CONTAINS it; nothing to push from that ref
|
|
288
|
+
|
|
289
|
+
git rev-list --left-right --count origin/main...HEAD # worktree-pre-push-scoped-pytest
|
|
290
|
+
-> 0 8 (behind 0, ahead 8)
|
|
291
|
+
git merge-base --is-ancestor HEAD origin/main
|
|
292
|
+
-> false NOT contained; the 8 are unique
|
|
293
|
+
```
|
|
294
|
+
|
|
295
|
+
Each of the 8 was additionally tested individually with
|
|
296
|
+
`git merge-base --is-ancestor <sha> origin/main`; none is in `origin/main`.
|
|
297
|
+
|
|
298
|
+
The audited session ref is not one I have worked on. The held work is on
|
|
299
|
+
`worktree-pre-push-scoped-pytest`. The push was NOT cancelled as stale,
|
|
300
|
+
because it does not target the audited ref -- but it remains UNPUSHED by
|
|
301
|
+
directive, pending explicit approval.
|
|
302
|
+
|
|
303
|
+
## 1c. origin/main is ONE JOB from CI-green
|
|
304
|
+
|
|
305
|
+
Receipts for `origin/main` = `dd692561`:
|
|
306
|
+
|
|
307
|
+
| Workflow | Conclusion | URL |
|
|
308
|
+
|---|---|---|
|
|
309
|
+
| Security Audit | success | actions/runs/30796434331 |
|
|
310
|
+
| SBOM (CycloneDX) | success | actions/runs/30794017044 |
|
|
311
|
+
| Bun Parity | success | actions/runs/30794017132 |
|
|
312
|
+
| Coverage (baseline) | success | actions/runs/30794017182 |
|
|
313
|
+
| pages build | success | actions/runs/30794016322 |
|
|
314
|
+
| **Tests** | **failure** | actions/runs/30794017151 |
|
|
315
|
+
|
|
316
|
+
The Tests failure is ONE job with ONE cause: `Shell tests (shard 3/4)` ->
|
|
317
|
+
`ShellCheck Linting FAILED`. Shards 0, 1 and 2 pass.
|
|
318
|
+
|
|
319
|
+
Measured on both sides of the fix:
|
|
320
|
+
|
|
321
|
+
- `origin/main`'s copies of the three files carry **7 shellcheck warning
|
|
322
|
+
sites** (test-iteration-grace 5, test-go-cargo-gate-timeout 1,
|
|
323
|
+
cleanup-test-processes 1).
|
|
324
|
+
- The held commit `a4d8b839` takes those same three files to **0**, and
|
|
325
|
+
`tests/run-shellcheck.sh` reports **360 passed / 0 failed**.
|
|
326
|
+
|
|
327
|
+
So the single blocker to a CI-green release candidate is already fixed and
|
|
328
|
+
held. No new work is required to clear it; only the push gate.
|
|
329
|
+
|
|
330
|
+
## 2. Held commits (8, unpushed)
|
|
331
|
+
|
|
332
|
+
| SHA | What it is |
|
|
333
|
+
|---|---|
|
|
334
|
+
| `a4d8b839` | Last red CI job: three PRE-EXISTING shellcheck warnings (not mine; last touched Jul 31 / Aug 1). One was a real hazard -- an unguarded `cd` that could make a gate examine the wrong tree and report a pass for a repo it never looked at. |
|
|
335
|
+
| `8c33d123` | Offline calibration audit. Discloses, before any number, that it measures agreement with the council majority and NOT accuracy. |
|
|
336
|
+
| `0f4da110` | T1/T2 benchmark designs frozen with held-out discipline. Unrun, unpaid. |
|
|
337
|
+
| `90468210` | Calibration wired into model-advisor as a ONE-WAY caveat. Pinned by test so it can never become a ranking input. |
|
|
338
|
+
| `19c2b5ad` | Per-command timeout on the completion-coverage probe. Root cause of the weekly audit's rc143, with the culprit named. |
|
|
339
|
+
| `fdba7023` | This handoff. |
|
|
340
|
+
| `12669cd3` | The no-spend outcome frontier: 3 held-out tasks, deterministic non-self-grading oracles, nothing run. |
|
|
341
|
+
| `ef0a6d49` | Spend-decision figures moved into the handoff so the decision needs one file, not two. |
|
|
342
|
+
|
|
343
|
+
## 3. Receipts
|
|
344
|
+
|
|
345
|
+
- **shellcheck: 360 passed, 0 failed.** Was 357/3.
|
|
346
|
+
- **Weekly-integrity rc143: root-caused and fixed.**
|
|
347
|
+
`test-completion-coverage.sh` probes ~274 command candidates as
|
|
348
|
+
`loki <c> --help`. `help` is among them (verified by re-running the
|
|
349
|
+
candidate extraction standalone), so it executed `loki help --help` -- the
|
|
350
|
+
unbounded self-delegation fixed in `ccf8dcbb`, which spawned a process per
|
|
351
|
+
level until the fork table was exhausted and the runner was SIGTERM'd.
|
|
352
|
+
Fixed with a PER-COMMAND 15s bound, never a blanket skip: skipping `help`
|
|
353
|
+
would have hidden the bug that needed fixing.
|
|
354
|
+
Proven non-vacuous by injecting a 60s hang into `doctor`:
|
|
355
|
+
`rc=1`, 34s, `PROBE HUNG: 'loki doctor --help' did not return within 15s`.
|
|
356
|
+
Without the hang: 6 passed, 0 failed, 20s.
|
|
357
|
+
- **CI on `dd692561` (last pushed SHA), terminal:** shards 0, 1, 2 GREEN --
|
|
358
|
+
the first shell shards to pass on a runner this session. Shard 3 red on
|
|
359
|
+
ShellCheck only, which `a4d8b839` fixes locally.
|
|
360
|
+
- **Local gate:** `scripts/local-ci.sh` exit 0, Failed 0, on every commit above.
|
|
361
|
+
- **Fork bomb:** contained 4,172 -> 375 processes. Root cause and counterfactual
|
|
362
|
+
in `ccf8dcbb` (already pushed).
|
|
363
|
+
|
|
364
|
+
## 4. The calibration audit has NEVER run on real data
|
|
365
|
+
|
|
366
|
+
Searched this machine for council transcripts:
|
|
367
|
+
|
|
368
|
+
```
|
|
369
|
+
find ~/git ~/loki-bench-v7 -type d -name transcripts -path "*council*" -> empty
|
|
370
|
+
find ~ -maxdepth 4 -type d -name council -> empty
|
|
371
|
+
```
|
|
372
|
+
|
|
373
|
+
**No real council transcripts exist here.** Every calibration number produced
|
|
374
|
+
so far is fixture-derived, and the fixtures were built to have a hand-checkable
|
|
375
|
+
answer rather than to resemble production. The audit is correct against
|
|
376
|
+
hand-derived arithmetic; it is unvalidated against reality.
|
|
377
|
+
|
|
378
|
+
This is stated because "the tool works" and "the tool has told us something
|
|
379
|
+
about our system" are different claims, and only the first is supported.
|
|
380
|
+
|
|
381
|
+
## 5. Risks
|
|
382
|
+
|
|
383
|
+
- **The calibration signal is not an accuracy signal, and it never becomes
|
|
384
|
+
one.** `tools/calibration-audit.py` scores agreement with the council
|
|
385
|
+
MAJORITY. The council outcome is derived from the votes, so a voter partly
|
|
386
|
+
causes its own label. A low Brier score there means "voted with the pack",
|
|
387
|
+
not "was correct". No artifact in this repo records whether the council was
|
|
388
|
+
right. This must stay a caveat; `90468210` pins by test that it cannot enter
|
|
389
|
+
model-advisor's machine-readable contract.
|
|
390
|
+
- **Local green is not CI green.** `a4d8b839` and `19c2b5ad` are verified only
|
|
391
|
+
on this machine. They are unpushed by instruction, so no runner has executed
|
|
392
|
+
them. Do not read the local 360/0 as a CI verdict.
|
|
393
|
+
- **Shard 3 has never been fully measured locally.** Its run was stopped during
|
|
394
|
+
the fork-bomb containment. Shards 0/1/2 were measured, partly during a window
|
|
395
|
+
when the process table was exhausted -- and in that state `grep` and `echo`
|
|
396
|
+
return EMPTY WITH NO ERROR, indistinguishable from "no matches". Readings
|
|
397
|
+
taken then are absent measurements, not results.
|
|
398
|
+
- **Three worktree agents' output was landed after review, and review found
|
|
399
|
+
real defects** (an unregistered guard that would never have run; a 48-vs-36
|
|
400
|
+
test-count misreport). Self-reports are not receipts.
|
|
401
|
+
- **Version skew is deliberate but load-bearing.** Anyone installing from npm
|
|
402
|
+
gets 9.8.1 and none of this work. If a user reports a bug fixed in 9.9-9.11,
|
|
403
|
+
that is why.
|
|
404
|
+
|
|
405
|
+
## 6. Founder-gated next action
|
|
406
|
+
|
|
407
|
+
Exactly one decision unblocks the rest. All five commits are gate-green
|
|
408
|
+
locally and held only by the standing directive.
|
|
409
|
+
|
|
410
|
+
**Option A -- release the gate.** Push the 5 commits. CI then runs them for the
|
|
411
|
+
first time, which is the only way to confirm shard 3 goes green and the
|
|
412
|
+
rc143 fix holds on a runner. No publish is implied by a push.
|
|
413
|
+
|
|
414
|
+
**Option B -- keep the gate, name the next target.** Work continues and
|
|
415
|
+
continues to be held. The next unblocked item is the outcome frontier design
|
|
416
|
+
(in flight, no-spend).
|
|
417
|
+
|
|
418
|
+
**Option C -- authorize spend.** Only this unlocks measuring OUTCOME quality
|
|
419
|
+
against competitors. Every axis in the scorecard that reads UNKNOWN reads that
|
|
420
|
+
way because no paid run has happened.
|
|
421
|
+
|
|
422
|
+
The design is ready and unrun: `docs/OUTCOME-FRONTIER.md`, three held-out
|
|
423
|
+
tasks with deterministic non-self-grading oracles, a frozen three-dimension
|
|
424
|
+
rubric, and seven dimensions deliberately left UNSCORED rather than
|
|
425
|
+
approximated with an LLM judge.
|
|
426
|
+
|
|
427
|
+
**Estimated cost: roughly $25 to $180 in provider spend and 4 to 12 hours of
|
|
428
|
+
wall clock**, for a 4-tool, 2-task matrix. That is an ESTIMATE with its
|
|
429
|
+
assumptions listed in the document, not a measurement, and the two runnable
|
|
430
|
+
tasks (brownfield, multi-file migration) are the ones priced. The scientific
|
|
431
|
+
task is PLAN ONLY -- it needs a paper chosen and its number transcribed before
|
|
432
|
+
it can run at all, and its results must never be reported as equally rigorous
|
|
433
|
+
to the other two.
|
|
434
|
+
|
|
435
|
+
What approving this buys: the first evidence about OUTCOME rather than CLI
|
|
436
|
+
affordances. What it does not buy: the seven unscored dimensions, which no
|
|
437
|
+
amount of spend makes deterministic.
|
|
438
|
+
|
|
439
|
+
Nothing in this handoff requires spend, and nothing in it has been published.
|
package/docs/INSTALLATION.md
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
The flagship product of [Autonomi](https://www.autonomi.dev/). Loki Mode is a spec-driven autonomous builder with a built-in trust layer that takes any spec to a deployed product and verifies completion with evidence (quality gates plus a completion council), not just a "done" claim. Complete installation instructions for all platforms and use cases.
|
|
4
4
|
|
|
5
|
-
**Version:**
|
|
5
|
+
**Version:** v9.8.1
|
|
6
6
|
|
|
7
7
|
---
|
|
8
8
|
|
|
@@ -36,7 +36,7 @@ Claude Code skill that compresses the model's OUTPUT tokens only (keeping all
|
|
|
36
36
|
technical substance). It activates on free-form generation (the main RARV dev
|
|
37
37
|
loop) and is HARD-SUPPRESSED on every trust-gate subcall (council votes, code
|
|
38
38
|
review verdict, evidence-related parses) so determinism is never affected.
|
|
39
|
-
- Claude-provider-only; runs are byte-identical on Codex /
|
|
39
|
+
- Claude-provider-only; runs are byte-identical on Cline / Codex / Aider / opencode.
|
|
40
40
|
- Default on; opt out with `LOKI_CAVEMAN=0`.
|
|
41
41
|
- Level: `LOKI_CAVEMAN_LEVEL` (default `full`; also `lite`, `ultra`, `wenyan*`).
|
|
42
42
|
- Pinned + vendor-less: `LOKI_CAVEMAN_VERSION` (default `1.9.0`); Loki bootstraps
|
|
@@ -46,7 +46,7 @@ review verdict, evidence-related parses) so determinism is never affected.
|
|
|
46
46
|
|
|
47
47
|
### Earlier highlights still in scope
|
|
48
48
|
- Bash-to-Bun runtime migration in progress (see `UPGRADING.md`)
|
|
49
|
-
- Provider-agnostic runtime: Claude (full),
|
|
49
|
+
- Provider-agnostic runtime: Claude (full), Cline, Codex, Aider, opencode (no vendor lock-in)
|
|
50
50
|
- Memory system (episodic / semantic / procedural)
|
|
51
51
|
- ChromaDB semantic code search via MCP
|
|
52
52
|
|
|
@@ -193,8 +193,9 @@ loki start ./spec.md
|
|
|
193
193
|
loki start ./spec.md --provider codex
|
|
194
194
|
loki start ./spec.md --provider cline
|
|
195
195
|
|
|
196
|
-
# Spec as a GitHub issue
|
|
197
|
-
loki start
|
|
196
|
+
# Spec as a GitHub issue -- pass it positionally, it is auto-detected
|
|
197
|
+
loki start https://github.com/owner/repo/issues/42
|
|
198
|
+
loki start owner/repo#42
|
|
198
199
|
|
|
199
200
|
# Spec as a YAML feature description
|
|
200
201
|
loki start ./feature.yaml
|
|
@@ -298,7 +299,7 @@ By default, prompt injection is **disabled** for enterprise safety:
|
|
|
298
299
|
./autonomy/run.sh ./my-spec.md
|
|
299
300
|
|
|
300
301
|
# Opt-in to enable prompt injection
|
|
301
|
-
|
|
302
|
+
LOKI_PROMPT_INJECTION=true ./autonomy/run.sh ./my-spec.md
|
|
302
303
|
```
|
|
303
304
|
|
|
304
305
|
#### Human Input Security
|
|
@@ -329,7 +330,13 @@ Loki Mode supports four active providers across three tiers, plus historical/upc
|
|
|
329
330
|
| `cline` | Active | Tier 2 | Full feature set; small models (<13B) may fail tool-use. |
|
|
330
331
|
| `codex` | Active | Tier 3 (degraded) | Sequential only, no Task tool; aligned with `@openai/codex` v0.125+. |
|
|
331
332
|
| `aider` | Active | Tier 3 (degraded) | Sequential only; `ollama_chat/<model>` works for local models. |
|
|
332
|
-
| `
|
|
333
|
+
| `opencode` | Active | Sequential | Model-agnostic route (`providers/opencode.sh`); autonomous flag `--auto`. Install: `npm install -g opencode-ai`. |
|
|
334
|
+
| `gemini` | REMOVED v7.5.18 | -- | Upstream Gemini CLI deprecated by Google. Runtime removed; `LOKI_PROVIDER=gemini` exits with a migration message. |
|
|
335
|
+
|
|
336
|
+
When `LOKI_PROVIDER` is unset, Loki auto-detects the first installed provider in
|
|
337
|
+
this order: `claude`, `cline`, `codex`, `aider`, `opencode`
|
|
338
|
+
(`providers/loader.sh:191`). An explicit `LOKI_PROVIDER` always wins and is
|
|
339
|
+
never silently substituted.
|
|
333
340
|
|
|
334
341
|
### Configuration
|
|
335
342
|
|
|
@@ -440,7 +447,7 @@ one-line export command and the compose volume to uncomment.
|
|
|
440
447
|
|
|
441
448
|
##### Other providers in Docker
|
|
442
449
|
|
|
443
|
-
The image ships only the Claude Code CLI. Codex,
|
|
450
|
+
The image ships only the Claude Code CLI. Cline, Codex, Aider, and opencode are
|
|
444
451
|
bring-your-own-CLI: install the provider CLI in a derived image (or mount it),
|
|
445
452
|
then select it with `-e LOKI_PROVIDER=<name>`. See
|
|
446
453
|
[DOCKER_README.md](../DOCKER_README.md) for details.
|
|
@@ -498,7 +505,7 @@ Claude Code skill that compresses the model's OUTPUT tokens only (keeping all
|
|
|
498
505
|
technical substance). It activates on free-form generation (the main RARV dev
|
|
499
506
|
loop) and is HARD-SUPPRESSED on every trust-gate subcall (council votes, code
|
|
500
507
|
review verdicts, evidence-related parses) so determinism is never affected. It
|
|
501
|
-
is Claude-provider-only; runs are byte-identical on Codex /
|
|
508
|
+
is Claude-provider-only; runs are byte-identical on Cline / Codex / Aider / opencode. These
|
|
502
509
|
variables are read in `autonomy/lib/claude-flags.sh`.
|
|
503
510
|
|
|
504
511
|
- `LOKI_CAVEMAN` (default on) -- set `LOKI_CAVEMAN=0` to disable the compressor.
|
|
@@ -843,7 +850,7 @@ The completion scripts support:
|
|
|
843
850
|
|
|
844
851
|
* **Smart Context**
|
|
845
852
|
|
|
846
|
-
* `loki start --provider <TAB>`
|
|
853
|
+
* `loki start --provider <TAB>` completes the supported providers (`claude`, `cline`, `codex`, `aider`, `opencode`; see `completions/loki.bash:24`).
|
|
847
854
|
* `loki start <TAB>` defaults to file completion for spec files (PRD templates, YAML).
|
|
848
855
|
|
|
849
856
|
* **Nested Commands**
|