loki-mode 9.12.5 → 9.16.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,254 @@
1
+ # Gate failure triage: exact-SHA classification
2
+
3
+ Ordered by the steer: reproduce locally against exact HEAD, classify each
4
+ failure as environment / baseline / candidate regression, evidence every claim
5
+ with a command and its output. Nothing here was pushed.
6
+
7
+ | Field | Value |
8
+ |---|---|
9
+ | HEAD at triage | `d230a3b2` (the steer named `09138e26`; that SHA is not in this worktree) |
10
+ | baseline compared | `dda8beec` = origin/main |
11
+ | held commits | `22199024`, `94315f35`, `d230a3b2`, plus `94f7639a` from this triage |
12
+ | pushed | **nothing** |
13
+
14
+ ## Classification
15
+
16
+ | Failure | Class | Evidence |
17
+ |---|---|---|
18
+ | `test-onboard-command` (6 of 9) | **BASELINE**, now FIXED | `autonomy/loki` byte-identical to origin/main; `git diff --name-only dda8beec..HEAD` returns zero matches for that path |
19
+ | `test-model-override` | **BASELINE**, 1 of 66, still open | identical failure at `dda8beec` and at HEAD: `Results: 65 passed, 1 failed (of 66)`, `EXIT=1` |
20
+ | `bun run typecheck` | **ENVIRONMENT** | `tsc` not installed locally; unchanged |
21
+
22
+ ## The onboard defect
23
+
24
+ `loki onboard --stdout` exited **141** and wrote **0 bytes** on this repo,
25
+ while passing on small fixtures.
26
+
27
+ ```
28
+ EXIT=141
29
+ STDOUT bytes: 0
30
+ STDERR bytes: 242
31
+ ```
32
+
33
+ 141 = 128+13 = SIGPIPE. `cmd_onboard`'s find fallback ends
34
+ `| sed | sort | head -200`. Under the file's `set -euo pipefail` (line 22),
35
+ `head` exits after 200 lines; `sort` -- which must consume ALL input before it
36
+ emits anything -- then takes SIGPIPE, the pipeline returns 141, and `-e` aborts
37
+ before a byte is written. The sibling `git ls-files` branch two lines above was
38
+ already guarded with `|| true`. The find branch never was.
39
+
40
+ `sort`, not `find`, is the process that dies. That distinction sets the test
41
+ size: see below.
42
+
43
+ ### Why no existing test caught it
44
+
45
+ **1. Every fixture was too small.** The residual output has to exceed the 64KB
46
+ pipe buffer before the signal lands. Measured:
47
+
48
+ | Fixture | Result vs unfixed code |
49
+ |---|---|
50
+ | 250 files | **passes** -- ~50 lines left after the cut, fits the buffer |
51
+ | 3000 files | **exit 141**, deterministic |
52
+
53
+ A 250-file regression test would have been worthless. I wrote one first,
54
+ confirmed it passed against the pre-fix binary, and resized it.
55
+
56
+ **2. This worktree never reached the guarded branch.** The check was
57
+ `[ -d "$target_path/.git" ]`, and in a git **worktree** `.git` is a pointer
58
+ **file**, not a directory:
59
+
60
+ ```
61
+ -rw-r--r-- 1 lokesh staff 83 Jul 31 19:25 .git
62
+ gitdir: /Users/lokesh/git/lokimode-anthropic/.git/worktrees/pre-push-scoped-pytest
63
+ ```
64
+
65
+ So every worktree silently fell through to the find path. Proven by trace:
66
+
67
+ ```
68
+ PRE-FIX ++ find ... -maxdepth 4
69
+ FIXED ++ git ls-files
70
+ ```
71
+
72
+ ### The fix
73
+
74
+ `-e` instead of `-d`, and `|| true` matching the sibling branch. Two lines.
75
+
76
+ ### Two sibling sites, quieter symptom
77
+
78
+ `cmd_explain` and `_docs_scan_project` carry the same pipeline at `head -500`.
79
+ Fixed alongside -- patching only the path the failure named would leave the
80
+ siblings broken.
81
+
82
+ They fail *differently*, which is why nothing ever caught them: both assign via
83
+ `local x=$(...)`, and `local` resets `$?`, swallowing the 141. Demonstrated:
84
+
85
+ ```
86
+ $ f() { local x=$(false | head -1); echo "rc=$?"; }
87
+ rc=0
88
+ ```
89
+
90
+ So they **silently truncate** their file tree instead of aborting. Same root
91
+ cause, no visible symptom.
92
+
93
+ ### Sweep
94
+
95
+ Three unguarded `sort | head -N` sites existed; zero remain. The sweep pattern
96
+ is not vacuous -- it matches 3 in the pre-fix file and 0 now.
97
+
98
+ The other two `-d .../.git` checks in the file (`loki:11113`, `loki:13473`) are
99
+ CORRECT as `-d`: one detects a clone (a worktree is not one), the other guards
100
+ `git init` on a fresh demo dir. Left alone.
101
+
102
+ ## Verification
103
+
104
+ | Check | Before | After |
105
+ |---|---|---|
106
+ | `test-onboard-command.sh` | 3/9 | **10/10** |
107
+ | new Test 10 vs pre-fix binary | **FAIL** (exit 141) | PASS |
108
+ | `test-onboard-json-injection-wave10.sh` | 2/2 | 2/2 |
109
+ | `test-contradiction-detection.sh` | 19/19 | 19/19 |
110
+ | `bash -n autonomy/loki` | OK | OK |
111
+
112
+ Test 10 was mutation-tested against `dda8beec`: it fails with the exact
113
+ diagnostic `exit 141 (SIGPIPE)` on the old code and passes on the new. It also
114
+ carries a vacuity guard rejecting exit 0 with under 100 bytes of output -- the
115
+ precise shape of the bug, since the abort produced exit 141 *and* silence.
116
+
117
+ ## Correction: I called test-model-override a non-failure before it finished
118
+
119
+ An earlier revision of THIS FILE classified `test-model-override` as "NOT A
120
+ FAILURE -- slow suite, mis-measured". That was wrong, and it was wrong in the
121
+ worst available way: I wrote the classification while the run was still
122
+ executing, from a partial log that showed 50 PASS and no failures yet.
123
+
124
+ The completed run:
125
+
126
+ ```
127
+ FAIL: architect no-cap mismatch: estimator='Opus' runner-dispatch='opus'
128
+ runner-tier='fable' (expected Opus,Sonnet / opus / fable)
129
+ Results: 65 passed, 1 failed (of 66)
130
+ EXIT=1
131
+ ```
132
+
133
+ The slowness was real -- a 224-cell parity matrix, each cell spawning a
134
+ `python3` that imports the FastAPI dashboard at ~0.46s -- and it was NOT the
135
+ explanation for the failure. Both things were true and I reported only the
136
+ convenient one.
137
+
138
+ The rule this violates is one already written down in this repo: an absent
139
+ measurement is not a measurement. A log with no FAIL line yet is not a log with
140
+ no failures; it is an unfinished log. I should have blocked on the EXIT marker
141
+ before classifying, exactly as I did for the onboard suite.
142
+
143
+ ### The actual failure
144
+
145
+ `tests/test-model-override.sh:827`. A three-way coherence assertion; two of the
146
+ three legs are correct:
147
+
148
+ | Leg | Expected | Actual |
149
+ |---|---|---|
150
+ | runner dispatch | `opus` | `opus` -- correct |
151
+ | runner tier (pre-collapse) | `fable` | `fable` -- correct |
152
+ | estimator quote | `Opus,Sonnet` | `Opus` -- **mismatch** |
153
+
154
+ So the runtime routing is right and only the cost QUOTE disagrees: the
155
+ estimator names one model where the run actually uses two. Per the test's own
156
+ comment (line 800), this is the known "estimator needs the sonnet5-default
157
+ update" case -- iter-1 collapses fable to opus, later iterations run the
158
+ development tier which defaults to sonnet since v7.104.0, so an honest quote
159
+ must name both.
160
+
161
+ It under-quotes cost. It does not mis-route a model.
162
+
163
+ ### Root cause: the fixture stopped producing enough iterations
164
+
165
+ The estimator is CORRECT. The expectation is only reachable when the estimate
166
+ spans more than one iteration, and the suite's fixture no longer does.
167
+
168
+ `autonomy/loki:18342` prices iteration 0 as Opus (the fable architect pass
169
+ collapsing to opus), and every LATER iteration through
170
+ `_priced_model_for(_dispatched_model)`, which defaults to Sonnet since
171
+ v7.104.0. So `Opus,Sonnet` requires **iterations >= 2**.
172
+
173
+ Measured on the suite's own fixture (`# PRD\nBuild a small todo API with one
174
+ endpoint.`, byte-identical to v7.104.0):
175
+
176
+ | Binary | tier | estimated iterations | nonzero models |
177
+ |---|---|---|---|
178
+ | `766219ac` (v7.104.0, where this was written and passed 66/0) | simple | **4** | `Opus,Sonnet` |
179
+ | HEAD | simple | **1** | `Opus` |
180
+
181
+ The complexity TIER is unchanged (`simple` in both). Only the iteration count
182
+ for that tier fell, 4 -> 1, and with a single iteration the loop never reaches
183
+ the branch where Sonnet appears.
184
+
185
+ Confirmed causal by holding the binary fixed and enlarging the input: a 24-
186
+ feature PRD at HEAD estimates 4 iterations and returns exactly `Opus,Sonnet`
187
+ (`{"Fable":0,"Opus":1,"Sonnet":3,"Haiku":0}`). Same code, more iterations,
188
+ expected answer.
189
+
190
+ So the assertion is a **stale coupling**: it encodes "a simple PRD takes
191
+ several iterations", which stopped being true. The v7.104.0 commit message
192
+ claims "locked by tests/test-model-override.sh (66/0)" -- that lock silently
193
+ came undone when the iteration estimate for simple PRDs changed.
194
+
195
+ ### Classification: BASELINE
196
+
197
+ Run against `dda8beec`'s `autonomy/loki` (my onboard fix reverted), the failure
198
+ is **identical**:
199
+
200
+ ```
201
+ FAIL: architect no-cap mismatch: estimator='Opus' ...
202
+ Results: 65 passed, 1 failed (of 66)
203
+ EXIT=1
204
+ ```
205
+
206
+ Not caused by any held commit. My `autonomy/loki` diff touches only three
207
+ tree-building sites and no pricing or routing code.
208
+
209
+ ### Not fixed here, deliberately
210
+
211
+ Two candidate fixes, and choosing between them is a product call I should not
212
+ make unilaterally:
213
+
214
+ 1. **Enlarge the fixture** so a multi-iteration estimate is exercised. Restores
215
+ the assertion's original intent (verify the architect collapse across a
216
+ real multi-iteration run) and keeps its coverage.
217
+ 2. **Weaken the assertion** to accept `Opus`. Cheaper, and wrong: it would
218
+ stop testing the later-iteration Sonnet attribution entirely.
219
+
220
+ (1) is almost certainly right, but it changes what the test measures, and the
221
+ prior instruction excluded runtime changes without an evidenced deterministic
222
+ requirement. The evidence is now here; the decision is not mine.
223
+
224
+ `docs/LOOP-CANDIDATE-PROPOSAL-v1.md` states the onboard defect "cannot be
225
+ addressed by any of [the six cheaper surfaces]" because it is a bash command
226
+ that never calls a model. That reasoning was right, and the conclusion drawn
227
+ from it was too weak: it is not a model-loop problem, it is a **two-line shell
228
+ bug**, and the correct action was to fix it rather than to route around it.
229
+
230
+ It also claimed the failure needed "a runtime fix to `autonomy/loki`" of
231
+ unknown size. Measured: 28 lines changed across three sites, all mechanical.
232
+
233
+ ## Gate status
234
+
235
+ Still not green, and this triage does not make it so.
236
+
237
+ - `test-onboard-command`: **RESOLVED** (3/9 -> 10/10)
238
+ - `test-model-override`: **1 of 66 still failing, BASELINE** -- a stale test
239
+ coupling, not an estimator defect. Root-caused (fixture no longer produces a
240
+ multi-iteration estimate); two candidate fixes named, neither applied.
241
+ - `bun run typecheck`: **unchanged**, `tsc` still absent
242
+
243
+ Two items remain, not one. Both are pre-existing at `dda8beec`; neither was
244
+ introduced by a held commit.
245
+
246
+ | Item | Needs |
247
+ |---|---|
248
+ | `bun run typecheck` | install the TS toolchain -- the environment fix already named |
249
+ | `test-model-override` | a decision between enlarging the fixture and weakening the assertion |
250
+
251
+ Neither is unblocked by the other, so "install tsc" was never sufficient on its
252
+ own. That was an error in the earlier proposal, which named a single
253
+ gate-closing action while a second real failure sat unclassified behind a
254
+ measurement I had cut short.
@@ -0,0 +1,167 @@
1
+ # loop-candidate-v1: proposal only
2
+
3
+ **GATE VERDICT: FALSE.** Nothing here is implemented, and nothing may be until
4
+ the gate closes. See the last section for the single gate-closing action.
5
+
6
+ | Gate | Verdict | Evidence |
7
+ |---|---|---|
8
+ | Local full tier | **DO NOT PUSH** | 161 passed, 3 failed, 0 deferred |
9
+ | Remote exact-SHA | **failure** | `Tests@dda8beec` completed/failure, Shell shard 2 |
10
+
11
+ ## 0. Why this proposal is small
12
+
13
+ The six cheaper surfaces come first by directive, and the honest finding is
14
+ that **the one evidenced product defect this session surfaced cannot be
15
+ addressed by any of them**: `loki onboard --stdout` emits 0 bytes and exits 0
16
+ on a 4100-file repository while working on a small one. It is a bash command
17
+ that never calls a model, so memory, retrieval, skills, prompts, tool
18
+ descriptions, compression and routing are all inapplicable by construction.
19
+
20
+ That leaves routing as the only surface with both an evidenced question and
21
+ existing instrumentation. This proposal is therefore scoped to **one routing
22
+ candidate**, and it is deliberately the smallest thing that could clear the
23
+ preregistered bar.
24
+
25
+ ## 1. Provenance
26
+
27
+ Every field pinned, no "current" or "latest":
28
+
29
+ | Field | Value |
30
+ |---|---|
31
+ | baseline runtime | `dda8beec` (origin/main), loki-mode 9.12.5 |
32
+ | candidate runtime | **none** -- this proposal changes no runtime code |
33
+ | harness | `tools/loop-harness-report.py` @ `2910a95d` (read-only) |
34
+ | task corpus | 1 spec, `wordcount` (pure helper + test), verbatim in both arms |
35
+ | grader | the produced test suite, executed; plus receipt `iterations.succeeded` |
36
+ | prompt version | unchanged; main-loop prompt is **NOT ATTRIBUTABLE** (assembled in memory, `run.sh:8987`) |
37
+ | skill version | unchanged |
38
+ | routing arm A | `LOKI_SESSION_MODEL=sonnet` |
39
+ | routing arm B | `LOKI_SESSION_MODEL=opus` |
40
+ | verifier set | unchanged; **records carry no cost/latency**, see §5 |
41
+
42
+ **The prompt row is a known provenance hole.** The main-loop prompt cannot be
43
+ versioned today, so any candidate that claims a prompt effect is unfalsifiable.
44
+ This proposal therefore claims none.
45
+
46
+ ## 2. Success and noise criteria, preregistered
47
+
48
+ - **Primary**: cost_usd per successful run, matched on the identical spec.
49
+ - **Secondary**: wall_clock_sec, and `progress duration_ms` reported separately
50
+ -- these are NOT the same quantity and are never summed. Wall clock includes
51
+ orchestration; progress duration is measured work.
52
+ - **Success requires**: `iterations.succeeded >= 1` AND the produced test suite
53
+ passes when executed. A receipt is not evidence the code works.
54
+ - **Noise rule**: a difference below **25%** on the primary is declared noise
55
+ and not reported as an effect. Justification: the two prior UNMATCHED runs
56
+ differed by 2.0x on the same nominal models, which bounds run-to-run
57
+ variance well above any effect a single pair could detect.
58
+ - **Minimum n**: 5 matched pairs. Below that, report "insufficient", never a
59
+ direction.
60
+
61
+ ## 3. Caps
62
+
63
+ | Cap | Value |
64
+ |---|---|
65
+ | retries | 0 -- a failed arm is recorded as failed, never retried |
66
+ | per-run timeout | 900s (matches the runs already executed) |
67
+ | iterations per run | `LOKI_MAX_ITERATIONS=2` |
68
+ | total spend ceiling | **$25**, hard stop; at ~$0.70/run that is ~35 runs |
69
+ | canary population | **none** -- offline only; no production traffic |
70
+ | canary window | n/a until offline clears |
71
+
72
+ ## 4. Evals
73
+
74
+ **Deterministic regression**: the existing suites, unchanged and frozen at the
75
+ baseline SHA. Any new failure disqualifies the candidate outright, regardless
76
+ of primary-metric movement.
77
+
78
+ **Ambitious artifact-level outcome**: one real `loki start` on a spec requiring
79
+ a multi-file artifact (module + test + usage doc), graded by (a) the produced
80
+ tests executing green, and (b) the receipt verifying via
81
+ `api_evidence.receipts_report`. Flat-or-better is required; a cost win with a
82
+ degraded artifact is a rejection, not a trade.
83
+
84
+ ## 5. Verifier lift versus p95 latency and cost -- BLOCKED, stated as such
85
+
86
+ This element **cannot be satisfied today** and the proposal does not pretend
87
+ otherwise.
88
+
89
+ Verifier records carry no cost, latency, criterion, or terminal-outcome effect:
90
+ `code_review_complete` emits exactly `review_id`, `source`, `iteration`; the
91
+ three gate functions (`_evidence_`, `_invariant_`, `_semantic_gate_and_surface`)
92
+ emit nothing structured at all. So the denominator for "lift vs p95 latency and
93
+ cost" does not exist, and no matched on/off cohort can be computed.
94
+
95
+ Closing this needs runtime instrumentation, which the directive excludes
96
+ without an evidenced deterministic requirement. **This proposal therefore makes
97
+ no verifier claim and proposes no verifier change.**
98
+
99
+ ## 6. One reversible candidate at a time
100
+
101
+ Exactly one variable moves: `LOKI_SESSION_MODEL`. It is an environment
102
+ variable, so reversal is unsetting it -- no code, no migration, no state.
103
+
104
+ Structured-trace learning is **read-only**: `loop-harness-report.py` reports
105
+ what traces contain and marks every underivable field UNKNOWN. It proposes
106
+ nothing automatically.
107
+
108
+ ## 7. Replay, rollback, retention
109
+
110
+ - **Frozen replay**: each run's receipt pins `base_sha`, `head_sha`,
111
+ `diff_sha256`, model, provider and cost. The spec is stored verbatim.
112
+ - **Automatic rollback**: none needed -- no runtime change to roll back. If the
113
+ candidate loses, the env var is simply not set.
114
+ - **Retained prior version**: baseline is `dda8beec` on origin/main, immutable.
115
+
116
+ ## 8. Trigger-to-receipt path
117
+
118
+ Already present and verified by execution, not by grep:
119
+
120
+ | Property | State |
121
+ |---|---|
122
+ | authentication | present |
123
+ | idempotency | present |
124
+ | dedupe | present -- `seen_delivery()`, lock-guarded bounded OrderedDict |
125
+ | backpressure | present -- bounded queue, 503 shed |
126
+ | bounded retry | present |
127
+ | timeout | present |
128
+ | dead-letter | **failures logged, not queryable** |
129
+
130
+ The one gap is a queryable failure record. It is narrow and **not proposed for
131
+ change** on this evidence.
132
+
133
+ ## 9. Human approval
134
+
135
+ Required for: promotion to default, any spend beyond the $25 ceiling,
136
+ destructive or security-boundary changes, and any external action (publish,
137
+ tag, post). Not required for: running the offline arms within the ceiling.
138
+
139
+ ## What would falsify this candidate
140
+
141
+ - any new deterministic regression -> reject
142
+ - artifact outcome degraded -> reject even if cheaper
143
+ - primary difference < 25% -> declare noise, no promotion
144
+ - fewer than 5 matched pairs -> report insufficient, no direction claimed
145
+
146
+ ## THE SINGLE GATE-CLOSING ACTION
147
+
148
+ The gate is red because of **three failures, none introduced by the held
149
+ commits**:
150
+
151
+ | Failure | Class | Fixable here? |
152
+ |---|---|---|
153
+ | `test-onboard-command` (6) | pre-existing; identical at `1c80c85ff~1` | needs a runtime fix to `autonomy/loki` |
154
+ | `test-model-override` (1) | pre-existing; 65/66 identical at `1c80c85ff~1` | unknown, undiagnosed |
155
+ | `bun run typecheck` | environmental; `tsc` not installed locally | one install |
156
+
157
+ **The single next action: install the TypeScript toolchain so
158
+ `bun run typecheck` can execute locally.** It is the only one of the three that
159
+ is a local environment gap rather than a code defect, it is non-destructive and
160
+ reversible, and it removes the one failure that is not telling us anything
161
+ about the repository.
162
+
163
+ That alone does not turn the gate green -- the two pre-existing suite failures
164
+ remain, and deciding whether to fix them, accept the remote gate as the
165
+ documented equivalent, or waive them is a founder call, not mine.
166
+
167
+ **Implementation stops here.**