loki-mode 9.12.6 → 9.16.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,167 @@
1
+ # loop-candidate-v1: proposal only
2
+
3
+ **GATE VERDICT: FALSE.** Nothing here is implemented, and nothing may be until
4
+ the gate closes. See the last section for the single gate-closing action.
5
+
6
+ | Gate | Verdict | Evidence |
7
+ |---|---|---|
8
+ | Local full tier | **DO NOT PUSH** | 161 passed, 3 failed, 0 deferred |
9
+ | Remote exact-SHA | **failure** | `Tests@dda8beec` completed/failure, Shell shard 2 |
10
+
11
+ ## 0. Why this proposal is small
12
+
13
+ The six cheaper surfaces come first by directive, and the honest finding is
14
+ that **the one evidenced product defect this session surfaced cannot be
15
+ addressed by any of them**: `loki onboard --stdout` emits 0 bytes and exits 0
16
+ on a 4100-file repository while working on a small one. It is a bash command
17
+ that never calls a model, so memory, retrieval, skills, prompts, tool
18
+ descriptions, compression and routing are all inapplicable by construction.
19
+
20
+ That leaves routing as the only surface with both an evidenced question and
21
+ existing instrumentation. This proposal is therefore scoped to **one routing
22
+ candidate**, and it is deliberately the smallest thing that could clear the
23
+ preregistered bar.
24
+
25
+ ## 1. Provenance
26
+
27
+ Every field pinned, no "current" or "latest":
28
+
29
+ | Field | Value |
30
+ |---|---|
31
+ | baseline runtime | `dda8beec` (origin/main), loki-mode 9.12.5 |
32
+ | candidate runtime | **none** -- this proposal changes no runtime code |
33
+ | harness | `tools/loop-harness-report.py` @ `2910a95d` (read-only) |
34
+ | task corpus | 1 spec, `wordcount` (pure helper + test), verbatim in both arms |
35
+ | grader | the produced test suite, executed; plus receipt `iterations.succeeded` |
36
+ | prompt version | unchanged; main-loop prompt is **NOT ATTRIBUTABLE** (assembled in memory, `run.sh:8987`) |
37
+ | skill version | unchanged |
38
+ | routing arm A | `LOKI_SESSION_MODEL=sonnet` |
39
+ | routing arm B | `LOKI_SESSION_MODEL=opus` |
40
+ | verifier set | unchanged; **records carry no cost/latency**, see §5 |
41
+
42
+ **The prompt row is a known provenance hole.** The main-loop prompt cannot be
43
+ versioned today, so any candidate that claims a prompt effect is unfalsifiable.
44
+ This proposal therefore claims none.
45
+
46
+ ## 2. Success and noise criteria, preregistered
47
+
48
+ - **Primary**: cost_usd per successful run, matched on the identical spec.
49
+ - **Secondary**: wall_clock_sec, and `progress duration_ms` reported separately
50
+ -- these are NOT the same quantity and are never summed. Wall clock includes
51
+ orchestration; progress duration is measured work.
52
+ - **Success requires**: `iterations.succeeded >= 1` AND the produced test suite
53
+ passes when executed. A receipt is not evidence the code works.
54
+ - **Noise rule**: a difference below **25%** on the primary is declared noise
55
+ and not reported as an effect. Justification: the two prior UNMATCHED runs
56
+ differed by 2.0x on the same nominal models, which bounds run-to-run
57
+ variance well above any effect a single pair could detect.
58
+ - **Minimum n**: 5 matched pairs. Below that, report "insufficient", never a
59
+ direction.
60
+
61
+ ## 3. Caps
62
+
63
+ | Cap | Value |
64
+ |---|---|
65
+ | retries | 0 -- a failed arm is recorded as failed, never retried |
66
+ | per-run timeout | 900s (matches the runs already executed) |
67
+ | iterations per run | `LOKI_MAX_ITERATIONS=2` |
68
+ | total spend ceiling | **$25**, hard stop; at ~$0.70/run that is ~35 runs |
69
+ | canary population | **none** -- offline only; no production traffic |
70
+ | canary window | n/a until offline clears |
71
+
72
+ ## 4. Evals
73
+
74
+ **Deterministic regression**: the existing suites, unchanged and frozen at the
75
+ baseline SHA. Any new failure disqualifies the candidate outright, regardless
76
+ of primary-metric movement.
77
+
78
+ **Ambitious artifact-level outcome**: one real `loki start` on a spec requiring
79
+ a multi-file artifact (module + test + usage doc), graded by (a) the produced
80
+ tests executing green, and (b) the receipt verifying via
81
+ `api_evidence.receipts_report`. Flat-or-better is required; a cost win with a
82
+ degraded artifact is a rejection, not a trade.
83
+
84
+ ## 5. Verifier lift versus p95 latency and cost -- BLOCKED, stated as such
85
+
86
+ This element **cannot be satisfied today** and the proposal does not pretend
87
+ otherwise.
88
+
89
+ Verifier records carry no cost, latency, criterion, or terminal-outcome effect:
90
+ `code_review_complete` emits exactly `review_id`, `source`, `iteration`; the
91
+ three gate functions (`_evidence_`, `_invariant_`, `_semantic_gate_and_surface`)
92
+ emit nothing structured at all. So the denominator for "lift vs p95 latency and
93
+ cost" does not exist, and no matched on/off cohort can be computed.
94
+
95
+ Closing this needs runtime instrumentation, which the directive excludes
96
+ without an evidenced deterministic requirement. **This proposal therefore makes
97
+ no verifier claim and proposes no verifier change.**
98
+
99
+ ## 6. One reversible candidate at a time
100
+
101
+ Exactly one variable moves: `LOKI_SESSION_MODEL`. It is an environment
102
+ variable, so reversal is unsetting it -- no code, no migration, no state.
103
+
104
+ Structured-trace learning is **read-only**: `loop-harness-report.py` reports
105
+ what traces contain and marks every underivable field UNKNOWN. It proposes
106
+ nothing automatically.
107
+
108
+ ## 7. Replay, rollback, retention
109
+
110
+ - **Frozen replay**: each run's receipt pins `base_sha`, `head_sha`,
111
+ `diff_sha256`, model, provider and cost. The spec is stored verbatim.
112
+ - **Automatic rollback**: none needed -- no runtime change to roll back. If the
113
+ candidate loses, the env var is simply not set.
114
+ - **Retained prior version**: baseline is `dda8beec` on origin/main, immutable.
115
+
116
+ ## 8. Trigger-to-receipt path
117
+
118
+ Already present and verified by execution, not by grep:
119
+
120
+ | Property | State |
121
+ |---|---|
122
+ | authentication | present |
123
+ | idempotency | present |
124
+ | dedupe | present -- `seen_delivery()`, lock-guarded bounded OrderedDict |
125
+ | backpressure | present -- bounded queue, 503 shed |
126
+ | bounded retry | present |
127
+ | timeout | present |
128
+ | dead-letter | **failures logged, not queryable** |
129
+
130
+ The one gap is a queryable failure record. It is narrow and **not proposed for
131
+ change** on this evidence.
132
+
133
+ ## 9. Human approval
134
+
135
+ Required for: promotion to default, any spend beyond the $25 ceiling,
136
+ destructive or security-boundary changes, and any external action (publish,
137
+ tag, post). Not required for: running the offline arms within the ceiling.
138
+
139
+ ## What would falsify this candidate
140
+
141
+ - any new deterministic regression -> reject
142
+ - artifact outcome degraded -> reject even if cheaper
143
+ - primary difference < 25% -> declare noise, no promotion
144
+ - fewer than 5 matched pairs -> report insufficient, no direction claimed
145
+
146
+ ## THE SINGLE GATE-CLOSING ACTION
147
+
148
+ The gate is red because of **three failures, none introduced by the held
149
+ commits**:
150
+
151
+ | Failure | Class | Fixable here? |
152
+ |---|---|---|
153
+ | `test-onboard-command` (6) | pre-existing; identical at `1c80c85ff~1` | needs a runtime fix to `autonomy/loki` |
154
+ | `test-model-override` (1) | pre-existing; 65/66 identical at `1c80c85ff~1` | unknown, undiagnosed |
155
+ | `bun run typecheck` | environmental; `tsc` not installed locally | one install |
156
+
157
+ **The single next action: install the TypeScript toolchain so
158
+ `bun run typecheck` can execute locally.** It is the only one of the three that
159
+ is a local environment gap rather than a code defect, it is non-destructive and
160
+ reversible, and it removes the one failure that is not telling us anything
161
+ about the repository.
162
+
163
+ That alone does not turn the gate green -- the two pre-existing suite failures
164
+ remain, and deciding whether to fix them, accept the remote gate as the
165
+ documented equivalent, or waive them is a founder call, not mine.
166
+
167
+ **Implementation stops here.**
@@ -629,3 +629,56 @@ Three procedural constraints were learned by getting them wrong first:
629
629
  2. run arms sequentially, or contention confounds the latency column
630
630
  3. verify the produced artifact by executing it -- a receipt records what a run
631
631
  claimed to do, not whether the code works
632
+
633
+ ## Read-only gap check: one evidenced product defect
634
+
635
+ The directive asks whether memory, skills, prompts, tool descriptions,
636
+ compression, or routing can address any evidenced gap. Running the full tier
637
+ surfaced a defect that none of those six can touch, because it is not a model
638
+ problem at all.
639
+
640
+ ### `loki onboard --stdout` produces nothing on a large repository
641
+
642
+ Reproduced deterministically, not inferred:
643
+
644
+ ```
645
+ small repo (1 source file), explicit path 394 bytes
646
+ small repo, cwd form 394 bytes
647
+ this repo (4100 tracked files), any form 0 bytes
648
+ this repo, --depth 1 0 bytes
649
+ ```
650
+
651
+ It is not the argument form: both spellings work on a small repo. It is not
652
+ scan depth: depth 1 fails identically. The command logs three INFO lines --
653
+ "Analyzing project at", "Depth: 2 | Format: markdown", "Scanning source
654
+ files..." -- then stops, writes no file, emits nothing to stdout, and **exits
655
+ 0**.
656
+
657
+ Exit 0 with no output is the worst available shape. A caller that checks the
658
+ exit code sees success; a caller that reads stdout gets an empty string and
659
+ cannot tell "this project has no structure" from "the command gave up". It is
660
+ the same class as the dashboard rendering an unmeasured cost as `$0.00`, which
661
+ this release line already fixed twice.
662
+
663
+ This is what `tests/test-onboard-command.sh` has been failing on: six
664
+ assertions, all downstream of empty output. The suite fails identically at
665
+ `1c80c85ff~1`, so it is pre-existing, not introduced this session.
666
+
667
+ ### Why none of the six named levers apply
668
+
669
+ Memory, retrieval, skills, prompts, tool descriptions, compression and routing
670
+ all shape what a MODEL is asked or told. This defect is in a bash analysis
671
+ command that never calls a model: it walks the filesystem and formats markdown.
672
+ No prompt change makes it emit output, and no routing decision is involved.
673
+
674
+ That is itself the useful finding. The directive's ordering -- try the cheap
675
+ model-facing levers before touching runtime -- is correct as a default and
676
+ does not apply here, and saying so is more honest than proposing a prompt
677
+ change that could not work.
678
+
679
+ ### Not fixed here
680
+
681
+ Fixing it means changing `autonomy/loki`, which is a runtime change this
682
+ directive excludes, and the failure mode (silent give-up at scale) needs its
683
+ own diagnosis before a patch. Recorded with a reproduction so it is actionable
684
+ rather than rediscovered.
@@ -0,0 +1,103 @@
1
+ # What our verification costs, and what it does not prove
2
+
3
+ This page exists because the most credible thing 8090 AI published was not a
4
+ capability claim. It was a cost:
5
+
6
+ > "the most common failure mode of vendor evaluation frameworks is to promise
7
+ > that no operational burden falls on the customer team. That promise is
8
+ > incompatible with a measurement signal that survives audit. The eval that
9
+ > costs nothing to run is, in our experience, the eval that cannot be defended
10
+ > in a regulatory inspection."
11
+
12
+ They name a seven-minute-per-document human cost, a 50-document golden dataset,
13
+ and several days per quarter of recalibration, and they refuse to automate the
14
+ one human signal away. That refusal is what makes the rest of their numbers
15
+ believable.
16
+
17
+ So here are ours. Every figure below was measured on this repository, and the
18
+ command that produces it is given so you can measure it yourself and get a
19
+ different answer on your hardware.
20
+
21
+ ## Measured cost
22
+
23
+ | Check | Cost | How it was measured |
24
+ |---|---|---|
25
+ | FULL local gate | 23 to 26 minutes | `LOCAL_CI_TIER=full bash scripts/local-ci.sh`, 166 checks, five runs on an M-series Mac |
26
+ | FAST local gate | about 1 minute | `bash scripts/local-ci.sh`, the documented pre-push tier |
27
+ | Shell suite, sharded | 352 seconds | 323 suites, `LOCAL_CI_SHARDS=4`; 1440 seconds serial, so 4.1x |
28
+ | `loki outcomes` | under 1 second per receipt | `git blame` and one `git log` per changed file |
29
+ | `loki intent status` | milliseconds | hash comparison against `.loki/spec/spec.lock` |
30
+ | Agent readiness | milliseconds | filesystem checks only, no network |
31
+ | Pre-edit snapshot | one `git diff` per run | write-once, at agent stop |
32
+
33
+ The FULL gate is the honest headline: **a full verification run costs about 25
34
+ minutes of wall clock**. We do not offer a mode that makes that free, because
35
+ the checks that take the time are the ones doing the work.
36
+
37
+ ## What our verification does NOT prove
38
+
39
+ Stated as plainly as we can, because a limits section that reads like marketing
40
+ is worse than none.
41
+
42
+ **A receipt is not proof the code is correct.** It proves a specific diff was
43
+ subjected to specific checks and records what each returned. A change can pass
44
+ every gate and still be wrong.
45
+
46
+ **The unsigned receipt path is forgeable.** Someone who rewrites both the facts
47
+ and the headline into a mutually consistent lie and recomputes the hash will
48
+ pass verification. That is defense in depth, not non-forgeability. Neutral
49
+ non-forgeability requires the signed path
50
+ (`LOKI_PROOF_GPG_KEY`, see [SIGNED-RECEIPTS.md](SIGNED-RECEIPTS.md)). We removed
51
+ our own "non-forgeable" claim in v7.111.0 after finding it false on that path.
52
+
53
+ **Only four of the eight quality gates are agent-independent.** Static analysis,
54
+ mock-integrity, test-mutation and documentation coverage do not ask a model
55
+ anything. The other four involve model judgment and are labelled ASSESSMENTS
56
+ rather than FACTS in every receipt.
57
+
58
+ **Verification cannot prove the spec was right.** This is the sharpest limit and
59
+ it is structural. Our gates prove code matches spec; if the spec diverged from
60
+ what you actually wanted, a passing gate is a correct answer to the wrong
61
+ question. `loki intent` measures that divergence where an intent has been
62
+ recorded, and reports UNKNOWN where it has not, which is most runs today.
63
+
64
+ **`loki outcomes` reports UNKNOWN on most existing receipts.** Measured on this
65
+ repository: 0 of 9 receipts are anchored, because 8 carry no recorded baseline
66
+ and 1 uses the empty-tree sha. A receipt is measured only when sha algebra
67
+ proves `base..head` is that change. We would rather print UNKNOWN than a
68
+ change-failure rate of 0.0 that no data supports.
69
+
70
+ **Generation is not air-gapped.** The verification path is local and offline;
71
+ generating code calls a model provider.
72
+
73
+ **We have no independent benchmark placement and no enterprise case studies.**
74
+ Neither exists yet. When they do they will be linked here, and until then their
75
+ absence is not evidence of anything except their absence.
76
+
77
+ ## What we refuse to build
78
+
79
+ Each of these would look good in a comparison table and would make the numbers
80
+ above less trustworthy:
81
+
82
+ - **A semantic fidelity score.** Asking a model whether a spec expresses an
83
+ intent and printing a percentage is a judgment wearing the costume of a
84
+ measurement.
85
+ - **A composite trust score.** Averaging a revert count, a hash comparison, a
86
+ path match and a model id yields a number whose movement nobody can explain.
87
+ - **A readiness percentage.** "There is no test command" tells you what to do.
88
+ "Readiness 62%" does not.
89
+ - **Any gate that punishes an agent for needing human edits.** It would train
90
+ the agent toward diffs nobody edits, which is not the same as good diffs.
91
+
92
+ ## Check any of this yourself
93
+
94
+ ```bash
95
+ LOCAL_CI_TIER=full bash scripts/local-ci.sh # the full gate, timed
96
+ loki outcomes --json # post-merge outcomes, or UNKNOWN with reasons
97
+ loki intent status --json # spec-vs-intent drift
98
+ loki proof verify <id> # re-hash a receipt, exit 1 on tamper
99
+ bash tests/test-competitor-verify-surface.sh # the competitor CLI measurement
100
+ ```
101
+
102
+ If a number here does not reproduce on your machine, that is a defect and we
103
+ want the report.
@@ -53,7 +53,7 @@ Measured on this machine, not asserted.
53
53
  | **1 Systems Thinking** | STRONG | 8 quality gates, RARV-C loop, council, Evidence Receipt, dual-route parity enforced by test |
54
54
  | **2 Speed** | MEASURED, UNOPTIMISED | agent call = **980s = 96%** of iteration; all gates together = 44s |
55
55
  | **3 Reliability** | STRONG, newly so | 73 mutation-proven trust invariants; four gates were shipping broken until v8.38.0 |
56
- | **4 Extensibility** | STRONG | 4 providers, 41 agent types, MCP (34 tools), plugin marketplace |
56
+ | **4 Extensibility** | STRONG | 4 providers, 41 agent types, MCP (36 tools), plugin marketplace |
57
57
  | **5 Feedback Loops / Evals** | **BROKEN** | `cost_usd == 0` on **3 of 5** efficiency records; no eval score for our own harness |
58
58
 
59
59
  **Principle 5 is the gap, and it is the one Wang weights highest.** He states the