loki-mode 9.12.5 → 9.16.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +81 -101
- package/SKILL.md +2 -2
- package/VERSION +1 -1
- package/autonomy/intent.sh +414 -0
- package/autonomy/issue-providers.sh +21 -0
- package/autonomy/lib/agent_readiness.py +202 -0
- package/autonomy/lib/claim_grounding.py +171 -0
- package/autonomy/lib/config-map.sh +10 -6
- package/autonomy/lib/decision_record.py +198 -0
- package/autonomy/lib/failure_memory.py +199 -0
- package/autonomy/lib/outcome_ledger.py +498 -0
- package/autonomy/lib/preedit_snapshot.py +216 -0
- package/autonomy/lib/verdict.py +204 -0
- package/autonomy/loki +378 -22
- package/autonomy/provider-offer.sh +25 -1
- package/autonomy/run.sh +516 -11
- package/autonomy/telemetry.sh +8 -1
- package/autonomy/verify.sh +10 -1
- package/completions/_loki +4 -0
- package/completions/loki.bash +2 -1
- package/dashboard/__init__.py +1 -1
- package/dashboard/api_operator.py +15 -2
- package/dashboard/api_v2.py +6 -1
- package/dashboard/control.py +62 -10
- package/dashboard/run.py +13 -2
- package/dashboard/scim.py +221 -0
- package/dashboard/server.py +48 -5
- package/docs/GATE-FAILURE-TRIAGE.md +254 -0
- package/docs/LOOP-CANDIDATE-PROPOSAL-v1.md +167 -0
- package/docs/LOOP-HARNESS-AUDIT.md +684 -0
- package/docs/VERIFICATION-COST.md +103 -0
- package/docs/WANG-PRINCIPLES-PLAN.md +1 -1
- package/loki-ts/dist/loki.js +402 -398
- package/mcp/__init__.py +1 -1
- package/mcp/_sdk_loader.py +25 -0
- package/package.json +1 -1
- package/plugins/loki-mode/.claude-plugin/plugin.json +1 -1
- package/tools/loop-harness-report.py +216 -0
|
@@ -0,0 +1,254 @@
|
|
|
1
|
+
# Gate failure triage: exact-SHA classification
|
|
2
|
+
|
|
3
|
+
Ordered by the steer: reproduce locally against exact HEAD, classify each
|
|
4
|
+
failure as environment / baseline / candidate regression, evidence every claim
|
|
5
|
+
with a command and its output. Nothing here was pushed.
|
|
6
|
+
|
|
7
|
+
| Field | Value |
|
|
8
|
+
|---|---|
|
|
9
|
+
| HEAD at triage | `d230a3b2` (the steer named `09138e26`; that SHA is not in this worktree) |
|
|
10
|
+
| baseline compared | `dda8beec` = origin/main |
|
|
11
|
+
| held commits | `22199024`, `94315f35`, `d230a3b2`, plus `94f7639a` from this triage |
|
|
12
|
+
| pushed | **nothing** |
|
|
13
|
+
|
|
14
|
+
## Classification
|
|
15
|
+
|
|
16
|
+
| Failure | Class | Evidence |
|
|
17
|
+
|---|---|---|
|
|
18
|
+
| `test-onboard-command` (6 of 9) | **BASELINE**, now FIXED | `autonomy/loki` byte-identical to origin/main; `git diff --name-only dda8beec..HEAD` returns zero matches for that path |
|
|
19
|
+
| `test-model-override` | **BASELINE**, 1 of 66, still open | identical failure at `dda8beec` and at HEAD: `Results: 65 passed, 1 failed (of 66)`, `EXIT=1` |
|
|
20
|
+
| `bun run typecheck` | **ENVIRONMENT** | `tsc` not installed locally; unchanged |
|
|
21
|
+
|
|
22
|
+
## The onboard defect
|
|
23
|
+
|
|
24
|
+
`loki onboard --stdout` exited **141** and wrote **0 bytes** on this repo,
|
|
25
|
+
while passing on small fixtures.
|
|
26
|
+
|
|
27
|
+
```
|
|
28
|
+
EXIT=141
|
|
29
|
+
STDOUT bytes: 0
|
|
30
|
+
STDERR bytes: 242
|
|
31
|
+
```
|
|
32
|
+
|
|
33
|
+
141 = 128+13 = SIGPIPE. `cmd_onboard`'s find fallback ends
|
|
34
|
+
`| sed | sort | head -200`. Under the file's `set -euo pipefail` (line 22),
|
|
35
|
+
`head` exits after 200 lines; `sort` -- which must consume ALL input before it
|
|
36
|
+
emits anything -- then takes SIGPIPE, the pipeline returns 141, and `-e` aborts
|
|
37
|
+
before a byte is written. The sibling `git ls-files` branch two lines above was
|
|
38
|
+
already guarded with `|| true`. The find branch never was.
|
|
39
|
+
|
|
40
|
+
`sort`, not `find`, is the process that dies. That distinction sets the test
|
|
41
|
+
size: see below.
|
|
42
|
+
|
|
43
|
+
### Why no existing test caught it
|
|
44
|
+
|
|
45
|
+
**1. Every fixture was too small.** The residual output has to exceed the 64KB
|
|
46
|
+
pipe buffer before the signal lands. Measured:
|
|
47
|
+
|
|
48
|
+
| Fixture | Result vs unfixed code |
|
|
49
|
+
|---|---|
|
|
50
|
+
| 250 files | **passes** -- ~50 lines left after the cut, fits the buffer |
|
|
51
|
+
| 3000 files | **exit 141**, deterministic |
|
|
52
|
+
|
|
53
|
+
A 250-file regression test would have been worthless. I wrote one first,
|
|
54
|
+
confirmed it passed against the pre-fix binary, and resized it.
|
|
55
|
+
|
|
56
|
+
**2. This worktree never reached the guarded branch.** The check was
|
|
57
|
+
`[ -d "$target_path/.git" ]`, and in a git **worktree** `.git` is a pointer
|
|
58
|
+
**file**, not a directory:
|
|
59
|
+
|
|
60
|
+
```
|
|
61
|
+
-rw-r--r-- 1 lokesh staff 83 Jul 31 19:25 .git
|
|
62
|
+
gitdir: /Users/lokesh/git/lokimode-anthropic/.git/worktrees/pre-push-scoped-pytest
|
|
63
|
+
```
|
|
64
|
+
|
|
65
|
+
So every worktree silently fell through to the find path. Proven by trace:
|
|
66
|
+
|
|
67
|
+
```
|
|
68
|
+
PRE-FIX ++ find ... -maxdepth 4
|
|
69
|
+
FIXED ++ git ls-files
|
|
70
|
+
```
|
|
71
|
+
|
|
72
|
+
### The fix
|
|
73
|
+
|
|
74
|
+
`-e` instead of `-d`, and `|| true` matching the sibling branch. Two lines.
|
|
75
|
+
|
|
76
|
+
### Two sibling sites, quieter symptom
|
|
77
|
+
|
|
78
|
+
`cmd_explain` and `_docs_scan_project` carry the same pipeline at `head -500`.
|
|
79
|
+
Fixed alongside -- patching only the path the failure named would leave the
|
|
80
|
+
siblings broken.
|
|
81
|
+
|
|
82
|
+
They fail *differently*, which is why nothing ever caught them: both assign via
|
|
83
|
+
`local x=$(...)`, and `local` resets `$?`, swallowing the 141. Demonstrated:
|
|
84
|
+
|
|
85
|
+
```
|
|
86
|
+
$ f() { local x=$(false | head -1); echo "rc=$?"; }
|
|
87
|
+
rc=0
|
|
88
|
+
```
|
|
89
|
+
|
|
90
|
+
So they **silently truncate** their file tree instead of aborting. Same root
|
|
91
|
+
cause, no visible symptom.
|
|
92
|
+
|
|
93
|
+
### Sweep
|
|
94
|
+
|
|
95
|
+
Three unguarded `sort | head -N` sites existed; zero remain. The sweep pattern
|
|
96
|
+
is not vacuous -- it matches 3 in the pre-fix file and 0 now.
|
|
97
|
+
|
|
98
|
+
The other two `-d .../.git` checks in the file (`loki:11113`, `loki:13473`) are
|
|
99
|
+
CORRECT as `-d`: one detects a clone (a worktree is not one), the other guards
|
|
100
|
+
`git init` on a fresh demo dir. Left alone.
|
|
101
|
+
|
|
102
|
+
## Verification
|
|
103
|
+
|
|
104
|
+
| Check | Before | After |
|
|
105
|
+
|---|---|---|
|
|
106
|
+
| `test-onboard-command.sh` | 3/9 | **10/10** |
|
|
107
|
+
| new Test 10 vs pre-fix binary | **FAIL** (exit 141) | PASS |
|
|
108
|
+
| `test-onboard-json-injection-wave10.sh` | 2/2 | 2/2 |
|
|
109
|
+
| `test-contradiction-detection.sh` | 19/19 | 19/19 |
|
|
110
|
+
| `bash -n autonomy/loki` | OK | OK |
|
|
111
|
+
|
|
112
|
+
Test 10 was mutation-tested against `dda8beec`: it fails with the exact
|
|
113
|
+
diagnostic `exit 141 (SIGPIPE)` on the old code and passes on the new. It also
|
|
114
|
+
carries a vacuity guard rejecting exit 0 with under 100 bytes of output -- the
|
|
115
|
+
precise shape of the bug, since the abort produced exit 141 *and* silence.
|
|
116
|
+
|
|
117
|
+
## Correction: I called test-model-override a non-failure before it finished
|
|
118
|
+
|
|
119
|
+
An earlier revision of THIS FILE classified `test-model-override` as "NOT A
|
|
120
|
+
FAILURE -- slow suite, mis-measured". That was wrong, and it was wrong in the
|
|
121
|
+
worst available way: I wrote the classification while the run was still
|
|
122
|
+
executing, from a partial log that showed 50 PASS and no failures yet.
|
|
123
|
+
|
|
124
|
+
The completed run:
|
|
125
|
+
|
|
126
|
+
```
|
|
127
|
+
FAIL: architect no-cap mismatch: estimator='Opus' runner-dispatch='opus'
|
|
128
|
+
runner-tier='fable' (expected Opus,Sonnet / opus / fable)
|
|
129
|
+
Results: 65 passed, 1 failed (of 66)
|
|
130
|
+
EXIT=1
|
|
131
|
+
```
|
|
132
|
+
|
|
133
|
+
The slowness was real -- a 224-cell parity matrix, each cell spawning a
|
|
134
|
+
`python3` that imports the FastAPI dashboard at ~0.46s -- and it was NOT the
|
|
135
|
+
explanation for the failure. Both things were true and I reported only the
|
|
136
|
+
convenient one.
|
|
137
|
+
|
|
138
|
+
The rule this violates is one already written down in this repo: an absent
|
|
139
|
+
measurement is not a measurement. A log with no FAIL line yet is not a log with
|
|
140
|
+
no failures; it is an unfinished log. I should have blocked on the EXIT marker
|
|
141
|
+
before classifying, exactly as I did for the onboard suite.
|
|
142
|
+
|
|
143
|
+
### The actual failure
|
|
144
|
+
|
|
145
|
+
`tests/test-model-override.sh:827`. A three-way coherence assertion; two of the
|
|
146
|
+
three legs are correct:
|
|
147
|
+
|
|
148
|
+
| Leg | Expected | Actual |
|
|
149
|
+
|---|---|---|
|
|
150
|
+
| runner dispatch | `opus` | `opus` -- correct |
|
|
151
|
+
| runner tier (pre-collapse) | `fable` | `fable` -- correct |
|
|
152
|
+
| estimator quote | `Opus,Sonnet` | `Opus` -- **mismatch** |
|
|
153
|
+
|
|
154
|
+
So the runtime routing is right and only the cost QUOTE disagrees: the
|
|
155
|
+
estimator names one model where the run actually uses two. Per the test's own
|
|
156
|
+
comment (line 800), this is the known "estimator needs the sonnet5-default
|
|
157
|
+
update" case -- iter-1 collapses fable to opus, later iterations run the
|
|
158
|
+
development tier which defaults to sonnet since v7.104.0, so an honest quote
|
|
159
|
+
must name both.
|
|
160
|
+
|
|
161
|
+
It under-quotes cost. It does not mis-route a model.
|
|
162
|
+
|
|
163
|
+
### Root cause: the fixture stopped producing enough iterations
|
|
164
|
+
|
|
165
|
+
The estimator is CORRECT. The expectation is only reachable when the estimate
|
|
166
|
+
spans more than one iteration, and the suite's fixture no longer does.
|
|
167
|
+
|
|
168
|
+
`autonomy/loki:18342` prices iteration 0 as Opus (the fable architect pass
|
|
169
|
+
collapsing to opus), and every LATER iteration through
|
|
170
|
+
`_priced_model_for(_dispatched_model)`, which defaults to Sonnet since
|
|
171
|
+
v7.104.0. So `Opus,Sonnet` requires **iterations >= 2**.
|
|
172
|
+
|
|
173
|
+
Measured on the suite's own fixture (`# PRD\nBuild a small todo API with one
|
|
174
|
+
endpoint.`, byte-identical to v7.104.0):
|
|
175
|
+
|
|
176
|
+
| Binary | tier | estimated iterations | nonzero models |
|
|
177
|
+
|---|---|---|---|
|
|
178
|
+
| `766219ac` (v7.104.0, where this was written and passed 66/0) | simple | **4** | `Opus,Sonnet` |
|
|
179
|
+
| HEAD | simple | **1** | `Opus` |
|
|
180
|
+
|
|
181
|
+
The complexity TIER is unchanged (`simple` in both). Only the iteration count
|
|
182
|
+
for that tier fell, 4 -> 1, and with a single iteration the loop never reaches
|
|
183
|
+
the branch where Sonnet appears.
|
|
184
|
+
|
|
185
|
+
Confirmed causal by holding the binary fixed and enlarging the input: a 24-
|
|
186
|
+
feature PRD at HEAD estimates 4 iterations and returns exactly `Opus,Sonnet`
|
|
187
|
+
(`{"Fable":0,"Opus":1,"Sonnet":3,"Haiku":0}`). Same code, more iterations,
|
|
188
|
+
expected answer.
|
|
189
|
+
|
|
190
|
+
So the assertion is a **stale coupling**: it encodes "a simple PRD takes
|
|
191
|
+
several iterations", which stopped being true. The v7.104.0 commit message
|
|
192
|
+
claims "locked by tests/test-model-override.sh (66/0)" -- that lock silently
|
|
193
|
+
came undone when the iteration estimate for simple PRDs changed.
|
|
194
|
+
|
|
195
|
+
### Classification: BASELINE
|
|
196
|
+
|
|
197
|
+
Run against `dda8beec`'s `autonomy/loki` (my onboard fix reverted), the failure
|
|
198
|
+
is **identical**:
|
|
199
|
+
|
|
200
|
+
```
|
|
201
|
+
FAIL: architect no-cap mismatch: estimator='Opus' ...
|
|
202
|
+
Results: 65 passed, 1 failed (of 66)
|
|
203
|
+
EXIT=1
|
|
204
|
+
```
|
|
205
|
+
|
|
206
|
+
Not caused by any held commit. My `autonomy/loki` diff touches only three
|
|
207
|
+
tree-building sites and no pricing or routing code.
|
|
208
|
+
|
|
209
|
+
### Not fixed here, deliberately
|
|
210
|
+
|
|
211
|
+
Two candidate fixes, and choosing between them is a product call I should not
|
|
212
|
+
make unilaterally:
|
|
213
|
+
|
|
214
|
+
1. **Enlarge the fixture** so a multi-iteration estimate is exercised. Restores
|
|
215
|
+
the assertion's original intent (verify the architect collapse across a
|
|
216
|
+
real multi-iteration run) and keeps its coverage.
|
|
217
|
+
2. **Weaken the assertion** to accept `Opus`. Cheaper, and wrong: it would
|
|
218
|
+
stop testing the later-iteration Sonnet attribution entirely.
|
|
219
|
+
|
|
220
|
+
(1) is almost certainly right, but it changes what the test measures, and the
|
|
221
|
+
prior instruction excluded runtime changes without an evidenced deterministic
|
|
222
|
+
requirement. The evidence is now here; the decision is not mine.
|
|
223
|
+
|
|
224
|
+
`docs/LOOP-CANDIDATE-PROPOSAL-v1.md` states the onboard defect "cannot be
|
|
225
|
+
addressed by any of [the six cheaper surfaces]" because it is a bash command
|
|
226
|
+
that never calls a model. That reasoning was right, and the conclusion drawn
|
|
227
|
+
from it was too weak: it is not a model-loop problem, it is a **two-line shell
|
|
228
|
+
bug**, and the correct action was to fix it rather than to route around it.
|
|
229
|
+
|
|
230
|
+
It also claimed the failure needed "a runtime fix to `autonomy/loki`" of
|
|
231
|
+
unknown size. Measured: 28 lines changed across three sites, all mechanical.
|
|
232
|
+
|
|
233
|
+
## Gate status
|
|
234
|
+
|
|
235
|
+
Still not green, and this triage does not make it so.
|
|
236
|
+
|
|
237
|
+
- `test-onboard-command`: **RESOLVED** (3/9 -> 10/10)
|
|
238
|
+
- `test-model-override`: **1 of 66 still failing, BASELINE** -- a stale test
|
|
239
|
+
coupling, not an estimator defect. Root-caused (fixture no longer produces a
|
|
240
|
+
multi-iteration estimate); two candidate fixes named, neither applied.
|
|
241
|
+
- `bun run typecheck`: **unchanged**, `tsc` still absent
|
|
242
|
+
|
|
243
|
+
Two items remain, not one. Both are pre-existing at `dda8beec`; neither was
|
|
244
|
+
introduced by a held commit.
|
|
245
|
+
|
|
246
|
+
| Item | Needs |
|
|
247
|
+
|---|---|
|
|
248
|
+
| `bun run typecheck` | install the TS toolchain -- the environment fix already named |
|
|
249
|
+
| `test-model-override` | a decision between enlarging the fixture and weakening the assertion |
|
|
250
|
+
|
|
251
|
+
Neither is unblocked by the other, so "install tsc" was never sufficient on its
|
|
252
|
+
own. That was an error in the earlier proposal, which named a single
|
|
253
|
+
gate-closing action while a second real failure sat unclassified behind a
|
|
254
|
+
measurement I had cut short.
|
|
@@ -0,0 +1,167 @@
|
|
|
1
|
+
# loop-candidate-v1: proposal only
|
|
2
|
+
|
|
3
|
+
**GATE VERDICT: FALSE.** Nothing here is implemented, and nothing may be until
|
|
4
|
+
the gate closes. See the last section for the single gate-closing action.
|
|
5
|
+
|
|
6
|
+
| Gate | Verdict | Evidence |
|
|
7
|
+
|---|---|---|
|
|
8
|
+
| Local full tier | **DO NOT PUSH** | 161 passed, 3 failed, 0 deferred |
|
|
9
|
+
| Remote exact-SHA | **failure** | `Tests@dda8beec` completed/failure, Shell shard 2 |
|
|
10
|
+
|
|
11
|
+
## 0. Why this proposal is small
|
|
12
|
+
|
|
13
|
+
The six cheaper surfaces come first by directive, and the honest finding is
|
|
14
|
+
that **the one evidenced product defect this session surfaced cannot be
|
|
15
|
+
addressed by any of them**: `loki onboard --stdout` emits 0 bytes and exits 0
|
|
16
|
+
on a 4100-file repository while working on a small one. It is a bash command
|
|
17
|
+
that never calls a model, so memory, retrieval, skills, prompts, tool
|
|
18
|
+
descriptions, compression and routing are all inapplicable by construction.
|
|
19
|
+
|
|
20
|
+
That leaves routing as the only surface with both an evidenced question and
|
|
21
|
+
existing instrumentation. This proposal is therefore scoped to **one routing
|
|
22
|
+
candidate**, and it is deliberately the smallest thing that could clear the
|
|
23
|
+
preregistered bar.
|
|
24
|
+
|
|
25
|
+
## 1. Provenance
|
|
26
|
+
|
|
27
|
+
Every field pinned, no "current" or "latest":
|
|
28
|
+
|
|
29
|
+
| Field | Value |
|
|
30
|
+
|---|---|
|
|
31
|
+
| baseline runtime | `dda8beec` (origin/main), loki-mode 9.12.5 |
|
|
32
|
+
| candidate runtime | **none** -- this proposal changes no runtime code |
|
|
33
|
+
| harness | `tools/loop-harness-report.py` @ `2910a95d` (read-only) |
|
|
34
|
+
| task corpus | 1 spec, `wordcount` (pure helper + test), verbatim in both arms |
|
|
35
|
+
| grader | the produced test suite, executed; plus receipt `iterations.succeeded` |
|
|
36
|
+
| prompt version | unchanged; main-loop prompt is **NOT ATTRIBUTABLE** (assembled in memory, `run.sh:8987`) |
|
|
37
|
+
| skill version | unchanged |
|
|
38
|
+
| routing arm A | `LOKI_SESSION_MODEL=sonnet` |
|
|
39
|
+
| routing arm B | `LOKI_SESSION_MODEL=opus` |
|
|
40
|
+
| verifier set | unchanged; **records carry no cost/latency**, see §5 |
|
|
41
|
+
|
|
42
|
+
**The prompt row is a known provenance hole.** The main-loop prompt cannot be
|
|
43
|
+
versioned today, so any candidate that claims a prompt effect is unfalsifiable.
|
|
44
|
+
This proposal therefore claims none.
|
|
45
|
+
|
|
46
|
+
## 2. Success and noise criteria, preregistered
|
|
47
|
+
|
|
48
|
+
- **Primary**: cost_usd per successful run, matched on the identical spec.
|
|
49
|
+
- **Secondary**: wall_clock_sec, and `progress duration_ms` reported separately
|
|
50
|
+
-- these are NOT the same quantity and are never summed. Wall clock includes
|
|
51
|
+
orchestration; progress duration is measured work.
|
|
52
|
+
- **Success requires**: `iterations.succeeded >= 1` AND the produced test suite
|
|
53
|
+
passes when executed. A receipt is not evidence the code works.
|
|
54
|
+
- **Noise rule**: a difference below **25%** on the primary is declared noise
|
|
55
|
+
and not reported as an effect. Justification: the two prior UNMATCHED runs
|
|
56
|
+
differed by 2.0x on the same nominal models, which bounds run-to-run
|
|
57
|
+
variance well above any effect a single pair could detect.
|
|
58
|
+
- **Minimum n**: 5 matched pairs. Below that, report "insufficient", never a
|
|
59
|
+
direction.
|
|
60
|
+
|
|
61
|
+
## 3. Caps
|
|
62
|
+
|
|
63
|
+
| Cap | Value |
|
|
64
|
+
|---|---|
|
|
65
|
+
| retries | 0 -- a failed arm is recorded as failed, never retried |
|
|
66
|
+
| per-run timeout | 900s (matches the runs already executed) |
|
|
67
|
+
| iterations per run | `LOKI_MAX_ITERATIONS=2` |
|
|
68
|
+
| total spend ceiling | **$25**, hard stop; at ~$0.70/run that is ~35 runs |
|
|
69
|
+
| canary population | **none** -- offline only; no production traffic |
|
|
70
|
+
| canary window | n/a until offline clears |
|
|
71
|
+
|
|
72
|
+
## 4. Evals
|
|
73
|
+
|
|
74
|
+
**Deterministic regression**: the existing suites, unchanged and frozen at the
|
|
75
|
+
baseline SHA. Any new failure disqualifies the candidate outright, regardless
|
|
76
|
+
of primary-metric movement.
|
|
77
|
+
|
|
78
|
+
**Ambitious artifact-level outcome**: one real `loki start` on a spec requiring
|
|
79
|
+
a multi-file artifact (module + test + usage doc), graded by (a) the produced
|
|
80
|
+
tests executing green, and (b) the receipt verifying via
|
|
81
|
+
`api_evidence.receipts_report`. Flat-or-better is required; a cost win with a
|
|
82
|
+
degraded artifact is a rejection, not a trade.
|
|
83
|
+
|
|
84
|
+
## 5. Verifier lift versus p95 latency and cost -- BLOCKED, stated as such
|
|
85
|
+
|
|
86
|
+
This element **cannot be satisfied today** and the proposal does not pretend
|
|
87
|
+
otherwise.
|
|
88
|
+
|
|
89
|
+
Verifier records carry no cost, latency, criterion, or terminal-outcome effect:
|
|
90
|
+
`code_review_complete` emits exactly `review_id`, `source`, `iteration`; the
|
|
91
|
+
three gate functions (`_evidence_`, `_invariant_`, `_semantic_gate_and_surface`)
|
|
92
|
+
emit nothing structured at all. So the denominator for "lift vs p95 latency and
|
|
93
|
+
cost" does not exist, and no matched on/off cohort can be computed.
|
|
94
|
+
|
|
95
|
+
Closing this needs runtime instrumentation, which the directive excludes
|
|
96
|
+
without an evidenced deterministic requirement. **This proposal therefore makes
|
|
97
|
+
no verifier claim and proposes no verifier change.**
|
|
98
|
+
|
|
99
|
+
## 6. One reversible candidate at a time
|
|
100
|
+
|
|
101
|
+
Exactly one variable moves: `LOKI_SESSION_MODEL`. It is an environment
|
|
102
|
+
variable, so reversal is unsetting it -- no code, no migration, no state.
|
|
103
|
+
|
|
104
|
+
Structured-trace learning is **read-only**: `loop-harness-report.py` reports
|
|
105
|
+
what traces contain and marks every underivable field UNKNOWN. It proposes
|
|
106
|
+
nothing automatically.
|
|
107
|
+
|
|
108
|
+
## 7. Replay, rollback, retention
|
|
109
|
+
|
|
110
|
+
- **Frozen replay**: each run's receipt pins `base_sha`, `head_sha`,
|
|
111
|
+
`diff_sha256`, model, provider and cost. The spec is stored verbatim.
|
|
112
|
+
- **Automatic rollback**: none needed -- no runtime change to roll back. If the
|
|
113
|
+
candidate loses, the env var is simply not set.
|
|
114
|
+
- **Retained prior version**: baseline is `dda8beec` on origin/main, immutable.
|
|
115
|
+
|
|
116
|
+
## 8. Trigger-to-receipt path
|
|
117
|
+
|
|
118
|
+
Already present and verified by execution, not by grep:
|
|
119
|
+
|
|
120
|
+
| Property | State |
|
|
121
|
+
|---|---|
|
|
122
|
+
| authentication | present |
|
|
123
|
+
| idempotency | present |
|
|
124
|
+
| dedupe | present -- `seen_delivery()`, lock-guarded bounded OrderedDict |
|
|
125
|
+
| backpressure | present -- bounded queue, 503 shed |
|
|
126
|
+
| bounded retry | present |
|
|
127
|
+
| timeout | present |
|
|
128
|
+
| dead-letter | **failures logged, not queryable** |
|
|
129
|
+
|
|
130
|
+
The one gap is a queryable failure record. It is narrow and **not proposed for
|
|
131
|
+
change** on this evidence.
|
|
132
|
+
|
|
133
|
+
## 9. Human approval
|
|
134
|
+
|
|
135
|
+
Required for: promotion to default, any spend beyond the $25 ceiling,
|
|
136
|
+
destructive or security-boundary changes, and any external action (publish,
|
|
137
|
+
tag, post). Not required for: running the offline arms within the ceiling.
|
|
138
|
+
|
|
139
|
+
## What would falsify this candidate
|
|
140
|
+
|
|
141
|
+
- any new deterministic regression -> reject
|
|
142
|
+
- artifact outcome degraded -> reject even if cheaper
|
|
143
|
+
- primary difference < 25% -> declare noise, no promotion
|
|
144
|
+
- fewer than 5 matched pairs -> report insufficient, no direction claimed
|
|
145
|
+
|
|
146
|
+
## THE SINGLE GATE-CLOSING ACTION
|
|
147
|
+
|
|
148
|
+
The gate is red because of **three failures, none introduced by the held
|
|
149
|
+
commits**:
|
|
150
|
+
|
|
151
|
+
| Failure | Class | Fixable here? |
|
|
152
|
+
|---|---|---|
|
|
153
|
+
| `test-onboard-command` (6) | pre-existing; identical at `1c80c85ff~1` | needs a runtime fix to `autonomy/loki` |
|
|
154
|
+
| `test-model-override` (1) | pre-existing; 65/66 identical at `1c80c85ff~1` | unknown, undiagnosed |
|
|
155
|
+
| `bun run typecheck` | environmental; `tsc` not installed locally | one install |
|
|
156
|
+
|
|
157
|
+
**The single next action: install the TypeScript toolchain so
|
|
158
|
+
`bun run typecheck` can execute locally.** It is the only one of the three that
|
|
159
|
+
is a local environment gap rather than a code defect, it is non-destructive and
|
|
160
|
+
reversible, and it removes the one failure that is not telling us anything
|
|
161
|
+
about the repository.
|
|
162
|
+
|
|
163
|
+
That alone does not turn the gate green -- the two pre-existing suite failures
|
|
164
|
+
remain, and deciding whether to fix them, accept the remote gate as the
|
|
165
|
+
documented equivalent, or waive them is a founder call, not mine.
|
|
166
|
+
|
|
167
|
+
**Implementation stops here.**
|