loki-mode 9.12.6 → 9.16.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +81 -101
- package/SKILL.md +2 -2
- package/VERSION +1 -1
- package/autonomy/intent.sh +414 -0
- package/autonomy/issue-providers.sh +21 -0
- package/autonomy/lib/agent_readiness.py +202 -0
- package/autonomy/lib/claim_grounding.py +171 -0
- package/autonomy/lib/config-map.sh +10 -6
- package/autonomy/lib/decision_record.py +198 -0
- package/autonomy/lib/failure_memory.py +199 -0
- package/autonomy/lib/outcome_ledger.py +498 -0
- package/autonomy/lib/preedit_snapshot.py +216 -0
- package/autonomy/lib/verdict.py +204 -0
- package/autonomy/loki +358 -14
- package/autonomy/provider-offer.sh +25 -1
- package/autonomy/run.sh +516 -11
- package/autonomy/telemetry.sh +8 -1
- package/completions/_loki +4 -0
- package/completions/loki.bash +2 -1
- package/dashboard/__init__.py +1 -1
- package/dashboard/run.py +13 -2
- package/dashboard/scim.py +221 -0
- package/dashboard/server.py +22 -0
- package/docs/GATE-FAILURE-TRIAGE.md +254 -0
- package/docs/LOOP-CANDIDATE-PROPOSAL-v1.md +167 -0
- package/docs/LOOP-HARNESS-AUDIT.md +53 -0
- package/docs/VERIFICATION-COST.md +103 -0
- package/docs/WANG-PRINCIPLES-PLAN.md +1 -1
- package/loki-ts/dist/loki.js +402 -398
- package/mcp/__init__.py +1 -1
- package/mcp/_sdk_loader.py +25 -0
- package/package.json +1 -1
- package/plugins/loki-mode/.claude-plugin/plugin.json +1 -1
|
@@ -0,0 +1,167 @@
|
|
|
1
|
+
# loop-candidate-v1: proposal only
|
|
2
|
+
|
|
3
|
+
**GATE VERDICT: FALSE.** Nothing here is implemented, and nothing may be until
|
|
4
|
+
the gate closes. See the last section for the single gate-closing action.
|
|
5
|
+
|
|
6
|
+
| Gate | Verdict | Evidence |
|
|
7
|
+
|---|---|---|
|
|
8
|
+
| Local full tier | **DO NOT PUSH** | 161 passed, 3 failed, 0 deferred |
|
|
9
|
+
| Remote exact-SHA | **failure** | `Tests@dda8beec` completed/failure, Shell shard 2 |
|
|
10
|
+
|
|
11
|
+
## 0. Why this proposal is small
|
|
12
|
+
|
|
13
|
+
The six cheaper surfaces come first by directive, and the honest finding is
|
|
14
|
+
that **the one evidenced product defect this session surfaced cannot be
|
|
15
|
+
addressed by any of them**: `loki onboard --stdout` emits 0 bytes and exits 0
|
|
16
|
+
on a 4100-file repository while working on a small one. It is a bash command
|
|
17
|
+
that never calls a model, so memory, retrieval, skills, prompts, tool
|
|
18
|
+
descriptions, compression and routing are all inapplicable by construction.
|
|
19
|
+
|
|
20
|
+
That leaves routing as the only surface with both an evidenced question and
|
|
21
|
+
existing instrumentation. This proposal is therefore scoped to **one routing
|
|
22
|
+
candidate**, and it is deliberately the smallest thing that could clear the
|
|
23
|
+
preregistered bar.
|
|
24
|
+
|
|
25
|
+
## 1. Provenance
|
|
26
|
+
|
|
27
|
+
Every field pinned, no "current" or "latest":
|
|
28
|
+
|
|
29
|
+
| Field | Value |
|
|
30
|
+
|---|---|
|
|
31
|
+
| baseline runtime | `dda8beec` (origin/main), loki-mode 9.12.5 |
|
|
32
|
+
| candidate runtime | **none** -- this proposal changes no runtime code |
|
|
33
|
+
| harness | `tools/loop-harness-report.py` @ `2910a95d` (read-only) |
|
|
34
|
+
| task corpus | 1 spec, `wordcount` (pure helper + test), verbatim in both arms |
|
|
35
|
+
| grader | the produced test suite, executed; plus receipt `iterations.succeeded` |
|
|
36
|
+
| prompt version | unchanged; main-loop prompt is **NOT ATTRIBUTABLE** (assembled in memory, `run.sh:8987`) |
|
|
37
|
+
| skill version | unchanged |
|
|
38
|
+
| routing arm A | `LOKI_SESSION_MODEL=sonnet` |
|
|
39
|
+
| routing arm B | `LOKI_SESSION_MODEL=opus` |
|
|
40
|
+
| verifier set | unchanged; **records carry no cost/latency**, see §5 |
|
|
41
|
+
|
|
42
|
+
**The prompt row is a known provenance hole.** The main-loop prompt cannot be
|
|
43
|
+
versioned today, so any candidate that claims a prompt effect is unfalsifiable.
|
|
44
|
+
This proposal therefore claims none.
|
|
45
|
+
|
|
46
|
+
## 2. Success and noise criteria, preregistered
|
|
47
|
+
|
|
48
|
+
- **Primary**: cost_usd per successful run, matched on the identical spec.
|
|
49
|
+
- **Secondary**: wall_clock_sec, and `progress duration_ms` reported separately
|
|
50
|
+
-- these are NOT the same quantity and are never summed. Wall clock includes
|
|
51
|
+
orchestration; progress duration is measured work.
|
|
52
|
+
- **Success requires**: `iterations.succeeded >= 1` AND the produced test suite
|
|
53
|
+
passes when executed. A receipt is not evidence the code works.
|
|
54
|
+
- **Noise rule**: a difference below **25%** on the primary is declared noise
|
|
55
|
+
and not reported as an effect. Justification: the two prior UNMATCHED runs
|
|
56
|
+
differed by 2.0x on the same nominal models, which bounds run-to-run
|
|
57
|
+
variance well above any effect a single pair could detect.
|
|
58
|
+
- **Minimum n**: 5 matched pairs. Below that, report "insufficient", never a
|
|
59
|
+
direction.
|
|
60
|
+
|
|
61
|
+
## 3. Caps
|
|
62
|
+
|
|
63
|
+
| Cap | Value |
|
|
64
|
+
|---|---|
|
|
65
|
+
| retries | 0 -- a failed arm is recorded as failed, never retried |
|
|
66
|
+
| per-run timeout | 900s (matches the runs already executed) |
|
|
67
|
+
| iterations per run | `LOKI_MAX_ITERATIONS=2` |
|
|
68
|
+
| total spend ceiling | **$25**, hard stop; at ~$0.70/run that is ~35 runs |
|
|
69
|
+
| canary population | **none** -- offline only; no production traffic |
|
|
70
|
+
| canary window | n/a until offline clears |
|
|
71
|
+
|
|
72
|
+
## 4. Evals
|
|
73
|
+
|
|
74
|
+
**Deterministic regression**: the existing suites, unchanged and frozen at the
|
|
75
|
+
baseline SHA. Any new failure disqualifies the candidate outright, regardless
|
|
76
|
+
of primary-metric movement.
|
|
77
|
+
|
|
78
|
+
**Ambitious artifact-level outcome**: one real `loki start` on a spec requiring
|
|
79
|
+
a multi-file artifact (module + test + usage doc), graded by (a) the produced
|
|
80
|
+
tests executing green, and (b) the receipt verifying via
|
|
81
|
+
`api_evidence.receipts_report`. Flat-or-better is required; a cost win with a
|
|
82
|
+
degraded artifact is a rejection, not a trade.
|
|
83
|
+
|
|
84
|
+
## 5. Verifier lift versus p95 latency and cost -- BLOCKED, stated as such
|
|
85
|
+
|
|
86
|
+
This element **cannot be satisfied today** and the proposal does not pretend
|
|
87
|
+
otherwise.
|
|
88
|
+
|
|
89
|
+
Verifier records carry no cost, latency, criterion, or terminal-outcome effect:
|
|
90
|
+
`code_review_complete` emits exactly `review_id`, `source`, `iteration`; the
|
|
91
|
+
three gate functions (`_evidence_`, `_invariant_`, `_semantic_gate_and_surface`)
|
|
92
|
+
emit nothing structured at all. So the denominator for "lift vs p95 latency and
|
|
93
|
+
cost" does not exist, and no matched on/off cohort can be computed.
|
|
94
|
+
|
|
95
|
+
Closing this needs runtime instrumentation, which the directive excludes
|
|
96
|
+
without an evidenced deterministic requirement. **This proposal therefore makes
|
|
97
|
+
no verifier claim and proposes no verifier change.**
|
|
98
|
+
|
|
99
|
+
## 6. One reversible candidate at a time
|
|
100
|
+
|
|
101
|
+
Exactly one variable moves: `LOKI_SESSION_MODEL`. It is an environment
|
|
102
|
+
variable, so reversal is unsetting it -- no code, no migration, no state.
|
|
103
|
+
|
|
104
|
+
Structured-trace learning is **read-only**: `loop-harness-report.py` reports
|
|
105
|
+
what traces contain and marks every underivable field UNKNOWN. It proposes
|
|
106
|
+
nothing automatically.
|
|
107
|
+
|
|
108
|
+
## 7. Replay, rollback, retention
|
|
109
|
+
|
|
110
|
+
- **Frozen replay**: each run's receipt pins `base_sha`, `head_sha`,
|
|
111
|
+
`diff_sha256`, model, provider and cost. The spec is stored verbatim.
|
|
112
|
+
- **Automatic rollback**: none needed -- no runtime change to roll back. If the
|
|
113
|
+
candidate loses, the env var is simply not set.
|
|
114
|
+
- **Retained prior version**: baseline is `dda8beec` on origin/main, immutable.
|
|
115
|
+
|
|
116
|
+
## 8. Trigger-to-receipt path
|
|
117
|
+
|
|
118
|
+
Already present and verified by execution, not by grep:
|
|
119
|
+
|
|
120
|
+
| Property | State |
|
|
121
|
+
|---|---|
|
|
122
|
+
| authentication | present |
|
|
123
|
+
| idempotency | present |
|
|
124
|
+
| dedupe | present -- `seen_delivery()`, lock-guarded bounded OrderedDict |
|
|
125
|
+
| backpressure | present -- bounded queue, 503 shed |
|
|
126
|
+
| bounded retry | present |
|
|
127
|
+
| timeout | present |
|
|
128
|
+
| dead-letter | **failures logged, not queryable** |
|
|
129
|
+
|
|
130
|
+
The one gap is a queryable failure record. It is narrow and **not proposed for
|
|
131
|
+
change** on this evidence.
|
|
132
|
+
|
|
133
|
+
## 9. Human approval
|
|
134
|
+
|
|
135
|
+
Required for: promotion to default, any spend beyond the $25 ceiling,
|
|
136
|
+
destructive or security-boundary changes, and any external action (publish,
|
|
137
|
+
tag, post). Not required for: running the offline arms within the ceiling.
|
|
138
|
+
|
|
139
|
+
## What would falsify this candidate
|
|
140
|
+
|
|
141
|
+
- any new deterministic regression -> reject
|
|
142
|
+
- artifact outcome degraded -> reject even if cheaper
|
|
143
|
+
- primary difference < 25% -> declare noise, no promotion
|
|
144
|
+
- fewer than 5 matched pairs -> report insufficient, no direction claimed
|
|
145
|
+
|
|
146
|
+
## THE SINGLE GATE-CLOSING ACTION
|
|
147
|
+
|
|
148
|
+
The gate is red because of **three failures, none introduced by the held
|
|
149
|
+
commits**:
|
|
150
|
+
|
|
151
|
+
| Failure | Class | Fixable here? |
|
|
152
|
+
|---|---|---|
|
|
153
|
+
| `test-onboard-command` (6) | pre-existing; identical at `1c80c85ff~1` | needs a runtime fix to `autonomy/loki` |
|
|
154
|
+
| `test-model-override` (1) | pre-existing; 65/66 identical at `1c80c85ff~1` | unknown, undiagnosed |
|
|
155
|
+
| `bun run typecheck` | environmental; `tsc` not installed locally | one install |
|
|
156
|
+
|
|
157
|
+
**The single next action: install the TypeScript toolchain so
|
|
158
|
+
`bun run typecheck` can execute locally.** It is the only one of the three that
|
|
159
|
+
is a local environment gap rather than a code defect, it is non-destructive and
|
|
160
|
+
reversible, and it removes the one failure that is not telling us anything
|
|
161
|
+
about the repository.
|
|
162
|
+
|
|
163
|
+
That alone does not turn the gate green -- the two pre-existing suite failures
|
|
164
|
+
remain, and deciding whether to fix them, accept the remote gate as the
|
|
165
|
+
documented equivalent, or waive them is a founder call, not mine.
|
|
166
|
+
|
|
167
|
+
**Implementation stops here.**
|
|
@@ -629,3 +629,56 @@ Three procedural constraints were learned by getting them wrong first:
|
|
|
629
629
|
2. run arms sequentially, or contention confounds the latency column
|
|
630
630
|
3. verify the produced artifact by executing it -- a receipt records what a run
|
|
631
631
|
claimed to do, not whether the code works
|
|
632
|
+
|
|
633
|
+
## Read-only gap check: one evidenced product defect
|
|
634
|
+
|
|
635
|
+
The directive asks whether memory, skills, prompts, tool descriptions,
|
|
636
|
+
compression, or routing can address any evidenced gap. Running the full tier
|
|
637
|
+
surfaced a defect that none of those six can touch, because it is not a model
|
|
638
|
+
problem at all.
|
|
639
|
+
|
|
640
|
+
### `loki onboard --stdout` produces nothing on a large repository
|
|
641
|
+
|
|
642
|
+
Reproduced deterministically, not inferred:
|
|
643
|
+
|
|
644
|
+
```
|
|
645
|
+
small repo (1 source file), explicit path 394 bytes
|
|
646
|
+
small repo, cwd form 394 bytes
|
|
647
|
+
this repo (4100 tracked files), any form 0 bytes
|
|
648
|
+
this repo, --depth 1 0 bytes
|
|
649
|
+
```
|
|
650
|
+
|
|
651
|
+
It is not the argument form: both spellings work on a small repo. It is not
|
|
652
|
+
scan depth: depth 1 fails identically. The command logs three INFO lines --
|
|
653
|
+
"Analyzing project at", "Depth: 2 | Format: markdown", "Scanning source
|
|
654
|
+
files..." -- then stops, writes no file, emits nothing to stdout, and **exits
|
|
655
|
+
0**.
|
|
656
|
+
|
|
657
|
+
Exit 0 with no output is the worst available shape. A caller that checks the
|
|
658
|
+
exit code sees success; a caller that reads stdout gets an empty string and
|
|
659
|
+
cannot tell "this project has no structure" from "the command gave up". It is
|
|
660
|
+
the same class as the dashboard rendering an unmeasured cost as `$0.00`, which
|
|
661
|
+
this release line already fixed twice.
|
|
662
|
+
|
|
663
|
+
This is what `tests/test-onboard-command.sh` has been failing on: six
|
|
664
|
+
assertions, all downstream of empty output. The suite fails identically at
|
|
665
|
+
`1c80c85ff~1`, so it is pre-existing, not introduced this session.
|
|
666
|
+
|
|
667
|
+
### Why none of the six named levers apply
|
|
668
|
+
|
|
669
|
+
Memory, retrieval, skills, prompts, tool descriptions, compression and routing
|
|
670
|
+
all shape what a MODEL is asked or told. This defect is in a bash analysis
|
|
671
|
+
command that never calls a model: it walks the filesystem and formats markdown.
|
|
672
|
+
No prompt change makes it emit output, and no routing decision is involved.
|
|
673
|
+
|
|
674
|
+
That is itself the useful finding. The directive's ordering -- try the cheap
|
|
675
|
+
model-facing levers before touching runtime -- is correct as a default and
|
|
676
|
+
does not apply here, and saying so is more honest than proposing a prompt
|
|
677
|
+
change that could not work.
|
|
678
|
+
|
|
679
|
+
### Not fixed here
|
|
680
|
+
|
|
681
|
+
Fixing it means changing `autonomy/loki`, which is a runtime change this
|
|
682
|
+
directive excludes, and the failure mode (silent give-up at scale) needs its
|
|
683
|
+
own diagnosis before a patch. Recorded with a reproduction so it is actionable
|
|
684
|
+
rather than rediscovered.
|
|
@@ -0,0 +1,103 @@
|
|
|
1
|
+
# What our verification costs, and what it does not prove
|
|
2
|
+
|
|
3
|
+
This page exists because the most credible thing 8090 AI published was not a
|
|
4
|
+
capability claim. It was a cost:
|
|
5
|
+
|
|
6
|
+
> "the most common failure mode of vendor evaluation frameworks is to promise
|
|
7
|
+
> that no operational burden falls on the customer team. That promise is
|
|
8
|
+
> incompatible with a measurement signal that survives audit. The eval that
|
|
9
|
+
> costs nothing to run is, in our experience, the eval that cannot be defended
|
|
10
|
+
> in a regulatory inspection."
|
|
11
|
+
|
|
12
|
+
They name a seven-minute-per-document human cost, a 50-document golden dataset,
|
|
13
|
+
and several days per quarter of recalibration, and they refuse to automate the
|
|
14
|
+
one human signal away. That refusal is what makes the rest of their numbers
|
|
15
|
+
believable.
|
|
16
|
+
|
|
17
|
+
So here are ours. Every figure below was measured on this repository, and the
|
|
18
|
+
command that produces it is given so you can measure it yourself and get a
|
|
19
|
+
different answer on your hardware.
|
|
20
|
+
|
|
21
|
+
## Measured cost
|
|
22
|
+
|
|
23
|
+
| Check | Cost | How it was measured |
|
|
24
|
+
|---|---|---|
|
|
25
|
+
| FULL local gate | 23 to 26 minutes | `LOCAL_CI_TIER=full bash scripts/local-ci.sh`, 166 checks, five runs on an M-series Mac |
|
|
26
|
+
| FAST local gate | about 1 minute | `bash scripts/local-ci.sh`, the documented pre-push tier |
|
|
27
|
+
| Shell suite, sharded | 352 seconds | 323 suites, `LOCAL_CI_SHARDS=4`; 1440 seconds serial, so 4.1x |
|
|
28
|
+
| `loki outcomes` | under 1 second per receipt | `git blame` and one `git log` per changed file |
|
|
29
|
+
| `loki intent status` | milliseconds | hash comparison against `.loki/spec/spec.lock` |
|
|
30
|
+
| Agent readiness | milliseconds | filesystem checks only, no network |
|
|
31
|
+
| Pre-edit snapshot | one `git diff` per run | write-once, at agent stop |
|
|
32
|
+
|
|
33
|
+
The FULL gate is the honest headline: **a full verification run costs about 25
|
|
34
|
+
minutes of wall clock**. We do not offer a mode that makes that free, because
|
|
35
|
+
the checks that take the time are the ones doing the work.
|
|
36
|
+
|
|
37
|
+
## What our verification does NOT prove
|
|
38
|
+
|
|
39
|
+
Stated as plainly as we can, because a limits section that reads like marketing
|
|
40
|
+
is worse than none.
|
|
41
|
+
|
|
42
|
+
**A receipt is not proof the code is correct.** It proves a specific diff was
|
|
43
|
+
subjected to specific checks and records what each returned. A change can pass
|
|
44
|
+
every gate and still be wrong.
|
|
45
|
+
|
|
46
|
+
**The unsigned receipt path is forgeable.** Someone who rewrites both the facts
|
|
47
|
+
and the headline into a mutually consistent lie and recomputes the hash will
|
|
48
|
+
pass verification. That is defense in depth, not non-forgeability. Neutral
|
|
49
|
+
non-forgeability requires the signed path
|
|
50
|
+
(`LOKI_PROOF_GPG_KEY`, see [SIGNED-RECEIPTS.md](SIGNED-RECEIPTS.md)). We removed
|
|
51
|
+
our own "non-forgeable" claim in v7.111.0 after finding it false on that path.
|
|
52
|
+
|
|
53
|
+
**Only four of the eight quality gates are agent-independent.** Static analysis,
|
|
54
|
+
mock-integrity, test-mutation and documentation coverage do not ask a model
|
|
55
|
+
anything. The other four involve model judgment and are labelled ASSESSMENTS
|
|
56
|
+
rather than FACTS in every receipt.
|
|
57
|
+
|
|
58
|
+
**Verification cannot prove the spec was right.** This is the sharpest limit and
|
|
59
|
+
it is structural. Our gates prove code matches spec; if the spec diverged from
|
|
60
|
+
what you actually wanted, a passing gate is a correct answer to the wrong
|
|
61
|
+
question. `loki intent` measures that divergence where an intent has been
|
|
62
|
+
recorded, and reports UNKNOWN where it has not, which is most runs today.
|
|
63
|
+
|
|
64
|
+
**`loki outcomes` reports UNKNOWN on most existing receipts.** Measured on this
|
|
65
|
+
repository: 0 of 9 receipts are anchored, because 8 carry no recorded baseline
|
|
66
|
+
and 1 uses the empty-tree sha. A receipt is measured only when sha algebra
|
|
67
|
+
proves `base..head` is that change. We would rather print UNKNOWN than a
|
|
68
|
+
change-failure rate of 0.0 that no data supports.
|
|
69
|
+
|
|
70
|
+
**Generation is not air-gapped.** The verification path is local and offline;
|
|
71
|
+
generating code calls a model provider.
|
|
72
|
+
|
|
73
|
+
**We have no independent benchmark placement and no enterprise case studies.**
|
|
74
|
+
Neither exists yet. When they do they will be linked here, and until then their
|
|
75
|
+
absence is not evidence of anything except their absence.
|
|
76
|
+
|
|
77
|
+
## What we refuse to build
|
|
78
|
+
|
|
79
|
+
Each of these would look good in a comparison table and would make the numbers
|
|
80
|
+
above less trustworthy:
|
|
81
|
+
|
|
82
|
+
- **A semantic fidelity score.** Asking a model whether a spec expresses an
|
|
83
|
+
intent and printing a percentage is a judgment wearing the costume of a
|
|
84
|
+
measurement.
|
|
85
|
+
- **A composite trust score.** Averaging a revert count, a hash comparison, a
|
|
86
|
+
path match and a model id yields a number whose movement nobody can explain.
|
|
87
|
+
- **A readiness percentage.** "There is no test command" tells you what to do.
|
|
88
|
+
"Readiness 62%" does not.
|
|
89
|
+
- **Any gate that punishes an agent for needing human edits.** It would train
|
|
90
|
+
the agent toward diffs nobody edits, which is not the same as good diffs.
|
|
91
|
+
|
|
92
|
+
## Check any of this yourself
|
|
93
|
+
|
|
94
|
+
```bash
|
|
95
|
+
LOCAL_CI_TIER=full bash scripts/local-ci.sh # the full gate, timed
|
|
96
|
+
loki outcomes --json # post-merge outcomes, or UNKNOWN with reasons
|
|
97
|
+
loki intent status --json # spec-vs-intent drift
|
|
98
|
+
loki proof verify <id> # re-hash a receipt, exit 1 on tamper
|
|
99
|
+
bash tests/test-competitor-verify-surface.sh # the competitor CLI measurement
|
|
100
|
+
```
|
|
101
|
+
|
|
102
|
+
If a number here does not reproduce on your machine, that is a defect and we
|
|
103
|
+
want the report.
|
|
@@ -53,7 +53,7 @@ Measured on this machine, not asserted.
|
|
|
53
53
|
| **1 Systems Thinking** | STRONG | 8 quality gates, RARV-C loop, council, Evidence Receipt, dual-route parity enforced by test |
|
|
54
54
|
| **2 Speed** | MEASURED, UNOPTIMISED | agent call = **980s = 96%** of iteration; all gates together = 44s |
|
|
55
55
|
| **3 Reliability** | STRONG, newly so | 73 mutation-proven trust invariants; four gates were shipping broken until v8.38.0 |
|
|
56
|
-
| **4 Extensibility** | STRONG | 4 providers, 41 agent types, MCP (
|
|
56
|
+
| **4 Extensibility** | STRONG | 4 providers, 41 agent types, MCP (36 tools), plugin marketplace |
|
|
57
57
|
| **5 Feedback Loops / Evals** | **BROKEN** | `cost_usd == 0` on **3 of 5** efficiency records; no eval score for our own harness |
|
|
58
58
|
|
|
59
59
|
**Principle 5 is the gap, and it is the one Wang weights highest.** He states the
|