loki-mode 9.12.6 → 9.16.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -194,7 +194,14 @@ _loki_known_command() {
194
194
  # a first run, never the user's paths, versions, or hostnames.
195
195
  _loki_known_blocker() {
196
196
  case "${1:-}" in
197
- no_provider|node|python3|jq|git|curl|disk|skill_symlink)
197
+ # not_logged_in is distinct from no_provider on purpose. They are the two
198
+ # halves of the same wall and they need opposite fixes: no_provider means
199
+ # "install something", not_logged_in means "authenticate the thing you
200
+ # already installed". Collapsing the second into `other` would make the
201
+ # single most common post-install failure unactionable in the data --
202
+ # and it is the worst-placed one, since without this the user only learns
203
+ # of it after confirming the spend.
204
+ no_provider|not_logged_in|node|python3|jq|git|curl|disk|skill_symlink)
198
205
  printf '%s' "$1" ;;
199
206
  *)
200
207
  printf 'other' ;;
package/completions/_loki CHANGED
@@ -171,6 +171,9 @@ function _loki_commands {
171
171
  'metrics:Session productivity report'
172
172
  'share:Share session report as GitHub Gist'
173
173
  'proof:Inspect/share proof-of-run artifacts'
174
+ 'outcomes:What happened to the work AFTER the receipt (reverted/reworked/survived)'
175
+ 'verdict:The five measured trust signals in one readable block'
176
+ 'readiness:Can an agent verify its own work in this repo? (measured, not LLM-scored)'
174
177
  'preview:Preview the locally-running app'
175
178
  'deploy:Deploy the built product (CI/CD-aware)'
176
179
  'context:Context window management'
@@ -190,6 +193,7 @@ function _loki_commands {
190
193
  'magic:Spec-driven component generation'
191
194
  'assets:Export/import shareable team assets'
192
195
  'spec:Living spec drift detection'
196
+ 'intent:Does the spec still say what was actually wanted'
193
197
  'verify:Deterministic PR verification'
194
198
  'grill:Interrogate a spec before building'
195
199
  'trust:Visible trust trajectory'
@@ -5,7 +5,8 @@ _loki_completion() {
5
5
  _init_completion || return
6
6
 
7
7
  # Main subcommands (must match autonomy/loki main case statement)
8
- local main_commands="start quick monitor demo tour welcome init stop pause resume steer status next ship dashboard web serve api sandbox notify import github issue config provider reset memory compound checkpoint council dogfood projects enterprise secrets cockpit secure own handoff doctor watchdog audit metrics syslog onboard share proof receipt explain plan report cost estimate kpis stats test ci watch telemetry agent context ctx code run export review optimize heal modernize migrate cluster worktree wt trigger failover remote deploy docker mcp magic assets analyze compliance crash open otel preview quickstart rc rollback self-update sentrux setup-skill spec state template trust trust-metrics ultracode update verify voice why wiki bench cleanup logs grill docs cp version completions help"
8
+ local main_commands="start quick monitor demo tour welcome init stop pause resume steer status next ship dashboard web serve api sandbox notify import github issue config provider reset memory compound checkpoint council dogfood projects enterprise secrets cockpit secure own handoff doctor watchdog audit metrics syslog onboard share proof receipt outcomes verdict readiness explain plan report cost estimate kpis stats test ci watch telemetry agent context ctx code run export review optimize heal modernize migrate cluster worktree wt trigger failover remote deploy docker mcp magic assets analyze compliance crash open otel preview quickstart rc rollback self-update sentrux setup-skill spec state template trust trust-metrics ultracode update verify voice why wiki bench cleanup logs grill docs cp version completions help"
9
+ local main_commands="start quick monitor demo tour welcome init stop pause resume steer status next ship dashboard web serve api sandbox notify import github issue intent config provider reset memory compound checkpoint council dogfood projects enterprise secrets cockpit secure own handoff doctor watchdog audit metrics syslog onboard share proof receipt explain plan report cost estimate kpis stats test ci watch telemetry agent context ctx code run export review optimize heal modernize migrate cluster worktree wt trigger failover remote deploy docker mcp magic assets analyze compliance crash open otel preview quickstart rc rollback self-update sentrux setup-skill spec state template trust trust-metrics ultracode update verify voice why wiki bench cleanup logs grill docs cp version completions help"
9
10
 
10
11
  # 1. If we are on the first argument (subcommand)
11
12
  if [[ $cword -eq 1 ]]; then
@@ -7,7 +7,7 @@ Modules:
7
7
  control: Session control API (start/stop/pause/resume)
8
8
  """
9
9
 
10
- __version__ = "9.12.6"
10
+ __version__ = "9.16.0"
11
11
 
12
12
  # Expose the control app for easy import
13
13
  try:
package/dashboard/run.py CHANGED
@@ -8,15 +8,26 @@ Usage:
8
8
  """
9
9
 
10
10
  import argparse
11
+ import os
11
12
  import sys
12
13
 
13
14
 
14
15
  def main():
15
16
  parser = argparse.ArgumentParser(description="Loki Mode Dashboard Server")
17
+ # Loopback by default, matching server.py:run_server and control.py. This
18
+ # launcher defaulted to 0.0.0.0, which published an UNAUTHENTICATED control
19
+ # plane on every interface: dashboard auth is opt-in
20
+ # (LOKI_ENTERPRISE_AUTH, default false in dashboard/auth.py), so with it
21
+ # unset -- the default -- anyone routable to the host could reach endpoints
22
+ # that stop builds and read receipts. A default must be safe on its own;
23
+ # exposing the network is a decision an operator makes explicitly, so
24
+ # --host 0.0.0.0 and LOKI_DASHBOARD_HOST both still work for the container
25
+ # and Helm paths that genuinely need it.
16
26
  parser.add_argument(
17
27
  "--host",
18
- default="0.0.0.0",
19
- help="Host to bind to (default: 0.0.0.0)",
28
+ default=os.environ.get("LOKI_DASHBOARD_HOST", "127.0.0.1"),
29
+ help="Host to bind to (default: $LOKI_DASHBOARD_HOST or 127.0.0.1). "
30
+ "Use 0.0.0.0 only with authentication enabled.",
20
31
  )
21
32
  parser.add_argument(
22
33
  "--port",
@@ -0,0 +1,221 @@
1
+ """SCIM 2.0 user provisioning: deprovisioning that actually revokes access.
2
+
3
+ WHY THIS EXISTS. Measured against our own code before writing any: OIDC is real
4
+ here (76 references in dashboard/auth.py) and SAML and SCIM are both ZERO. That
5
+ is the whole enterprise identity gap -- not authentication as a category, which
6
+ we have, but PROVISIONING, which we do not.
7
+
8
+ It matters more than its size suggests. With OIDC alone, an employee removed
9
+ from the IdP stops being able to LOG IN, but any token already issued to them
10
+ keeps working until it expires. For a buyer whose security review asks "show me
11
+ that a terminated employee loses access within minutes", that is the answer that
12
+ fails the review. SCIM is the protocol that answers it.
13
+
14
+ WORTH KNOWING ABOUT THE COMPETITION. Factory ships SCIM. Devin's SCIM is
15
+ documented only for Devin DESKTOP -- the page instructs you to configure it at
16
+ windsurf.com/team/settings under the app name "Windsurf" -- and no page documents
17
+ SCIM provisioning for Devin CLOUD sessions, which get SAML/OIDC plus IdP-group
18
+ sync only. So this is table stakes against one competitor and a genuine gap in
19
+ the other.
20
+
21
+ WHAT THIS IS NOT. It is not a second identity system. It maps IdP group
22
+ membership onto the scope hierarchy dashboard/auth.py already enforces
23
+ (* -> control -> write -> read), so there is exactly one authorization model and
24
+ SCIM only decides who holds which scope. A parallel permission model would
25
+ eventually disagree with the real one, and the disagreement would be invisible
26
+ until an audit.
27
+
28
+ DEPROVISIONING IS SOFT BY DEFAULT, AND THAT IS DELIBERATE. `active: false`
29
+ revokes access immediately but PRESERVES the record, because an audit trail that
30
+ loses the identity of whoever performed past actions is not an audit trail. Hard
31
+ delete is available and separate, so destroying history is always an explicit act.
32
+ """
33
+
34
+ from __future__ import annotations
35
+
36
+ import json
37
+ import os
38
+ from datetime import datetime, timezone
39
+ from typing import Optional
40
+
41
+ SCIM_SCHEMA_USER = "urn:ietf:params:scim:schemas:core:2.0:User"
42
+ SCIM_SCHEMA_LIST = "urn:ietf:params:scim:api:messages:2.0:ListResponse"
43
+ SCIM_SCHEMA_ERROR = "urn:ietf:params:scim:api:messages:2.0:Error"
44
+
45
+ # IdP group -> our scope. Mirrors the hierarchy in dashboard/auth.py rather than
46
+ # inventing a second one: * grants everything, control grants write and read.
47
+ # An unrecognised group maps to "read", never to a higher scope -- an unknown
48
+ # group is not evidence of privilege, and defaulting upward is how a
49
+ # misconfigured IdP silently grants admin.
50
+ GROUP_SCOPE_MAP = {
51
+ "loki-admins": "*",
52
+ "loki-owners": "*",
53
+ "loki-operators": "control",
54
+ "loki-writers": "write",
55
+ "loki-readers": "read",
56
+ }
57
+ DEFAULT_SCOPE = "read"
58
+
59
+
60
+ def _now():
61
+ return datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
62
+
63
+
64
+ def _store_path(loki_dir):
65
+ return os.path.join(loki_dir, "scim", "users.json")
66
+
67
+
68
+ def _load(loki_dir):
69
+ p = _store_path(loki_dir)
70
+ if not os.path.isfile(p):
71
+ return {}
72
+ try:
73
+ with open(p, "r", encoding="utf-8") as fh:
74
+ return json.load(fh)
75
+ except (OSError, ValueError):
76
+ # A corrupt store must not silently become an empty one: returning {}
77
+ # here would read as "no users provisioned" and could re-grant access to
78
+ # someone who was deprovisioned. Signalled to the caller instead.
79
+ return None
80
+
81
+
82
+ def _save(loki_dir, data):
83
+ p = _store_path(loki_dir)
84
+ os.makedirs(os.path.dirname(p), exist_ok=True)
85
+ tmp = p + ".tmp"
86
+ with open(tmp, "w", encoding="utf-8") as fh:
87
+ json.dump(data, fh, indent=2, sort_keys=True)
88
+ fh.write("\n")
89
+ # Atomic replace: a half-written user store is an access-control decision
90
+ # made from corrupt data.
91
+ os.replace(tmp, p)
92
+
93
+
94
+ def scope_for_groups(groups):
95
+ """Highest scope any of the user's groups grants.
96
+
97
+ Highest-wins, because a user in both loki-readers and loki-admins is an
98
+ admin. Unknown groups contribute DEFAULT_SCOPE and never more.
99
+ """
100
+ rank = {"read": 1, "write": 2, "control": 3, "*": 4}
101
+ best = DEFAULT_SCOPE
102
+ for g in groups or []:
103
+ s = GROUP_SCOPE_MAP.get(str(g).strip().lower(), DEFAULT_SCOPE)
104
+ if rank.get(s, 1) > rank.get(best, 1):
105
+ best = s
106
+ return best
107
+
108
+
109
+ def provision(loki_dir, user_name, external_id=None, groups=None, active=True):
110
+ """Create or update a user. Returns a SCIM User resource."""
111
+ if not user_name:
112
+ return {"schemas": [SCIM_SCHEMA_ERROR], "status": "400",
113
+ "detail": "userName is required"}
114
+
115
+ data = _load(loki_dir)
116
+ if data is None:
117
+ return {"schemas": [SCIM_SCHEMA_ERROR], "status": "500",
118
+ "detail": "user store is unreadable; refusing to make an access "
119
+ "decision from corrupt data"}
120
+
121
+ uid = external_id or user_name
122
+ existing = data.get(uid, {})
123
+ rec = {
124
+ "id": uid,
125
+ "userName": user_name,
126
+ "externalId": external_id or "",
127
+ "groups": list(groups or []),
128
+ "scope": scope_for_groups(groups),
129
+ "active": bool(active),
130
+ "created": existing.get("created") or _now(),
131
+ "lastModified": _now(),
132
+ }
133
+ data[uid] = rec
134
+ _save(loki_dir, data)
135
+ return to_scim_user(rec)
136
+
137
+
138
+ def deprovision(loki_dir, uid, hard=False):
139
+ """Revoke access. Soft by default so the audit trail survives.
140
+
141
+ An audit log that cannot say WHO performed an action is not an audit log, so
142
+ a departed user's record is retained by default and merely marked inactive.
143
+ Hard delete exists and is separate, so destroying history is always explicit.
144
+ """
145
+ data = _load(loki_dir)
146
+ if data is None:
147
+ return {"schemas": [SCIM_SCHEMA_ERROR], "status": "500",
148
+ "detail": "user store is unreadable"}
149
+ if uid not in data:
150
+ return {"schemas": [SCIM_SCHEMA_ERROR], "status": "404",
151
+ "detail": f"no such user: {uid}"}
152
+
153
+ if hard:
154
+ data.pop(uid)
155
+ _save(loki_dir, data)
156
+ return {"status": "204", "deleted": uid, "audit_trail_preserved": False}
157
+
158
+ data[uid]["active"] = False
159
+ data[uid]["lastModified"] = _now()
160
+ _save(loki_dir, data)
161
+ return {"status": "200", "deactivated": uid, "audit_trail_preserved": True}
162
+
163
+
164
+ def is_active(loki_dir, uid):
165
+ """Authorization answer for one user. Fails CLOSED on any doubt.
166
+
167
+ An unreadable store, an unknown user, and an inactive user all return False.
168
+ A provisioning system that fails open is worse than none, because it creates
169
+ the belief that deprovisioning works.
170
+ """
171
+ data = _load(loki_dir)
172
+ if data is None:
173
+ return False
174
+ rec = data.get(uid)
175
+ if not rec:
176
+ return False
177
+ return bool(rec.get("active"))
178
+
179
+
180
+ def effective_scope(loki_dir, uid):
181
+ """The scope a user actually holds, or None when they hold none."""
182
+ data = _load(loki_dir)
183
+ if data is None:
184
+ return None
185
+ rec = data.get(uid)
186
+ if not rec or not rec.get("active"):
187
+ return None
188
+ return rec.get("scope") or DEFAULT_SCOPE
189
+
190
+
191
+ def to_scim_user(rec):
192
+ return {
193
+ "schemas": [SCIM_SCHEMA_USER],
194
+ "id": rec["id"],
195
+ "userName": rec["userName"],
196
+ "externalId": rec.get("externalId", ""),
197
+ "active": rec.get("active", False),
198
+ "groups": [{"value": g} for g in rec.get("groups", [])],
199
+ "meta": {
200
+ "resourceType": "User",
201
+ "created": rec.get("created"),
202
+ "lastModified": rec.get("lastModified"),
203
+ },
204
+ # Not part of the SCIM spec, surfaced because an operator reading this
205
+ # response needs to see what the group mapping actually granted rather
206
+ # than deriving it.
207
+ "lokiScope": rec.get("scope"),
208
+ }
209
+
210
+
211
+ def list_users(loki_dir, active_only=False):
212
+ data = _load(loki_dir)
213
+ if data is None:
214
+ return {"schemas": [SCIM_SCHEMA_ERROR], "status": "500",
215
+ "detail": "user store is unreadable"}
216
+ recs = [r for r in data.values() if (not active_only or r.get("active"))]
217
+ return {
218
+ "schemas": [SCIM_SCHEMA_LIST],
219
+ "totalResults": len(recs),
220
+ "Resources": [to_scim_user(r) for r in sorted(recs, key=lambda r: r["id"])],
221
+ }
@@ -1224,6 +1224,28 @@ class WebSocketBoundaryMiddleware:
1224
1224
  # accept-then-drop, which reads to a client as a flaky network.
1225
1225
  await send({"type": "websocket.close", "code": 1008})
1226
1226
  return
1227
+ # Cross-Site WebSocket Hijacking. The address check above proves the
1228
+ # PEER is local; it says nothing about who told that peer to connect.
1229
+ # A browser on any origin can open ws://127.0.0.1:57374/ws -- the
1230
+ # request comes FROM the loopback interface and passes the check
1231
+ # above, and CORS does not apply to WebSocket upgrades, so no
1232
+ # existing control stops it. /ws/collab is writable, so this is not
1233
+ # read-only exposure.
1234
+ #
1235
+ # Same-origin browsers send Origin; non-browser clients (CLI, tests,
1236
+ # health probes) send none. An ABSENT Origin is therefore allowed --
1237
+ # it cannot be forged by a page, since browsers always attach it --
1238
+ # while a PRESENT one must be in the allowlist the HTTP path already
1239
+ # uses. That keeps every scripted client working and closes the
1240
+ # browser-driven path, without a second source of truth for origins.
1241
+ origin = None
1242
+ for raw_name, raw_value in scope.get("headers", []):
1243
+ if raw_name == b"origin":
1244
+ origin = raw_value.decode("latin-1").strip()
1245
+ break
1246
+ if origin and "*" not in _cors_origins and origin not in _cors_origins:
1247
+ await send({"type": "websocket.close", "code": 1008})
1248
+ return
1227
1249
  await self.app(scope, receive, send)
1228
1250
 
1229
1251
 
@@ -0,0 +1,254 @@
1
+ # Gate failure triage: exact-SHA classification
2
+
3
+ Ordered by the steer: reproduce locally against exact HEAD, classify each
4
+ failure as environment / baseline / candidate regression, evidence every claim
5
+ with a command and its output. Nothing here was pushed.
6
+
7
+ | Field | Value |
8
+ |---|---|
9
+ | HEAD at triage | `d230a3b2` (the steer named `09138e26`; that SHA is not in this worktree) |
10
+ | baseline compared | `dda8beec` = origin/main |
11
+ | held commits | `22199024`, `94315f35`, `d230a3b2`, plus `94f7639a` from this triage |
12
+ | pushed | **nothing** |
13
+
14
+ ## Classification
15
+
16
+ | Failure | Class | Evidence |
17
+ |---|---|---|
18
+ | `test-onboard-command` (6 of 9) | **BASELINE**, now FIXED | `autonomy/loki` byte-identical to origin/main; `git diff --name-only dda8beec..HEAD` returns zero matches for that path |
19
+ | `test-model-override` | **BASELINE**, 1 of 66, still open | identical failure at `dda8beec` and at HEAD: `Results: 65 passed, 1 failed (of 66)`, `EXIT=1` |
20
+ | `bun run typecheck` | **ENVIRONMENT** | `tsc` not installed locally; unchanged |
21
+
22
+ ## The onboard defect
23
+
24
+ `loki onboard --stdout` exited **141** and wrote **0 bytes** on this repo,
25
+ while passing on small fixtures.
26
+
27
+ ```
28
+ EXIT=141
29
+ STDOUT bytes: 0
30
+ STDERR bytes: 242
31
+ ```
32
+
33
+ 141 = 128+13 = SIGPIPE. `cmd_onboard`'s find fallback ends
34
+ `| sed | sort | head -200`. Under the file's `set -euo pipefail` (line 22),
35
+ `head` exits after 200 lines; `sort` -- which must consume ALL input before it
36
+ emits anything -- then takes SIGPIPE, the pipeline returns 141, and `-e` aborts
37
+ before a byte is written. The sibling `git ls-files` branch two lines above was
38
+ already guarded with `|| true`. The find branch never was.
39
+
40
+ `sort`, not `find`, is the process that dies. That distinction sets the test
41
+ size: see below.
42
+
43
+ ### Why no existing test caught it
44
+
45
+ **1. Every fixture was too small.** The residual output has to exceed the 64KB
46
+ pipe buffer before the signal lands. Measured:
47
+
48
+ | Fixture | Result vs unfixed code |
49
+ |---|---|
50
+ | 250 files | **passes** -- ~50 lines left after the cut, fits the buffer |
51
+ | 3000 files | **exit 141**, deterministic |
52
+
53
+ A 250-file regression test would have been worthless. I wrote one first,
54
+ confirmed it passed against the pre-fix binary, and resized it.
55
+
56
+ **2. This worktree never reached the guarded branch.** The check was
57
+ `[ -d "$target_path/.git" ]`, and in a git **worktree** `.git` is a pointer
58
+ **file**, not a directory:
59
+
60
+ ```
61
+ -rw-r--r-- 1 lokesh staff 83 Jul 31 19:25 .git
62
+ gitdir: /Users/lokesh/git/lokimode-anthropic/.git/worktrees/pre-push-scoped-pytest
63
+ ```
64
+
65
+ So every worktree silently fell through to the find path. Proven by trace:
66
+
67
+ ```
68
+ PRE-FIX ++ find ... -maxdepth 4
69
+ FIXED ++ git ls-files
70
+ ```
71
+
72
+ ### The fix
73
+
74
+ `-e` instead of `-d`, and `|| true` matching the sibling branch. Two lines.
75
+
76
+ ### Two sibling sites, quieter symptom
77
+
78
+ `cmd_explain` and `_docs_scan_project` carry the same pipeline at `head -500`.
79
+ Fixed alongside -- patching only the path the failure named would leave the
80
+ siblings broken.
81
+
82
+ They fail *differently*, which is why nothing ever caught them: both assign via
83
+ `local x=$(...)`, and `local` resets `$?`, swallowing the 141. Demonstrated:
84
+
85
+ ```
86
+ $ f() { local x=$(false | head -1); echo "rc=$?"; }
87
+ rc=0
88
+ ```
89
+
90
+ So they **silently truncate** their file tree instead of aborting. Same root
91
+ cause, no visible symptom.
92
+
93
+ ### Sweep
94
+
95
+ Three unguarded `sort | head -N` sites existed; zero remain. The sweep pattern
96
+ is not vacuous -- it matches 3 in the pre-fix file and 0 now.
97
+
98
+ The other two `-d .../.git` checks in the file (`loki:11113`, `loki:13473`) are
99
+ CORRECT as `-d`: one detects a clone (a worktree is not one), the other guards
100
+ `git init` on a fresh demo dir. Left alone.
101
+
102
+ ## Verification
103
+
104
+ | Check | Before | After |
105
+ |---|---|---|
106
+ | `test-onboard-command.sh` | 3/9 | **10/10** |
107
+ | new Test 10 vs pre-fix binary | **FAIL** (exit 141) | PASS |
108
+ | `test-onboard-json-injection-wave10.sh` | 2/2 | 2/2 |
109
+ | `test-contradiction-detection.sh` | 19/19 | 19/19 |
110
+ | `bash -n autonomy/loki` | OK | OK |
111
+
112
+ Test 10 was mutation-tested against `dda8beec`: it fails with the exact
113
+ diagnostic `exit 141 (SIGPIPE)` on the old code and passes on the new. It also
114
+ carries a vacuity guard rejecting exit 0 with under 100 bytes of output -- the
115
+ precise shape of the bug, since the abort produced exit 141 *and* silence.
116
+
117
+ ## Correction: I called test-model-override a non-failure before it finished
118
+
119
+ An earlier revision of THIS FILE classified `test-model-override` as "NOT A
120
+ FAILURE -- slow suite, mis-measured". That was wrong, and it was wrong in the
121
+ worst available way: I wrote the classification while the run was still
122
+ executing, from a partial log that showed 50 PASS and no failures yet.
123
+
124
+ The completed run:
125
+
126
+ ```
127
+ FAIL: architect no-cap mismatch: estimator='Opus' runner-dispatch='opus'
128
+ runner-tier='fable' (expected Opus,Sonnet / opus / fable)
129
+ Results: 65 passed, 1 failed (of 66)
130
+ EXIT=1
131
+ ```
132
+
133
+ The slowness was real -- a 224-cell parity matrix, each cell spawning a
134
+ `python3` that imports the FastAPI dashboard at ~0.46s -- and it was NOT the
135
+ explanation for the failure. Both things were true and I reported only the
136
+ convenient one.
137
+
138
+ The rule this violates is one already written down in this repo: an absent
139
+ measurement is not a measurement. A log with no FAIL line yet is not a log with
140
+ no failures; it is an unfinished log. I should have blocked on the EXIT marker
141
+ before classifying, exactly as I did for the onboard suite.
142
+
143
+ ### The actual failure
144
+
145
+ `tests/test-model-override.sh:827`. A three-way coherence assertion; two of the
146
+ three legs are correct:
147
+
148
+ | Leg | Expected | Actual |
149
+ |---|---|---|
150
+ | runner dispatch | `opus` | `opus` -- correct |
151
+ | runner tier (pre-collapse) | `fable` | `fable` -- correct |
152
+ | estimator quote | `Opus,Sonnet` | `Opus` -- **mismatch** |
153
+
154
+ So the runtime routing is right and only the cost QUOTE disagrees: the
155
+ estimator names one model where the run actually uses two. Per the test's own
156
+ comment (line 800), this is the known "estimator needs the sonnet5-default
157
+ update" case -- iter-1 collapses fable to opus, later iterations run the
158
+ development tier which defaults to sonnet since v7.104.0, so an honest quote
159
+ must name both.
160
+
161
+ It under-quotes cost. It does not mis-route a model.
162
+
163
+ ### Root cause: the fixture stopped producing enough iterations
164
+
165
+ The estimator is CORRECT. The expectation is only reachable when the estimate
166
+ spans more than one iteration, and the suite's fixture no longer does.
167
+
168
+ `autonomy/loki:18342` prices iteration 0 as Opus (the fable architect pass
169
+ collapsing to opus), and every LATER iteration through
170
+ `_priced_model_for(_dispatched_model)`, which defaults to Sonnet since
171
+ v7.104.0. So `Opus,Sonnet` requires **iterations >= 2**.
172
+
173
+ Measured on the suite's own fixture (`# PRD\nBuild a small todo API with one
174
+ endpoint.`, byte-identical to v7.104.0):
175
+
176
+ | Binary | tier | estimated iterations | nonzero models |
177
+ |---|---|---|---|
178
+ | `766219ac` (v7.104.0, where this was written and passed 66/0) | simple | **4** | `Opus,Sonnet` |
179
+ | HEAD | simple | **1** | `Opus` |
180
+
181
+ The complexity TIER is unchanged (`simple` in both). Only the iteration count
182
+ for that tier fell, 4 -> 1, and with a single iteration the loop never reaches
183
+ the branch where Sonnet appears.
184
+
185
+ Confirmed causal by holding the binary fixed and enlarging the input: a 24-
186
+ feature PRD at HEAD estimates 4 iterations and returns exactly `Opus,Sonnet`
187
+ (`{"Fable":0,"Opus":1,"Sonnet":3,"Haiku":0}`). Same code, more iterations,
188
+ expected answer.
189
+
190
+ So the assertion is a **stale coupling**: it encodes "a simple PRD takes
191
+ several iterations", which stopped being true. The v7.104.0 commit message
192
+ claims "locked by tests/test-model-override.sh (66/0)" -- that lock silently
193
+ came undone when the iteration estimate for simple PRDs changed.
194
+
195
+ ### Classification: BASELINE
196
+
197
+ Run against `dda8beec`'s `autonomy/loki` (my onboard fix reverted), the failure
198
+ is **identical**:
199
+
200
+ ```
201
+ FAIL: architect no-cap mismatch: estimator='Opus' ...
202
+ Results: 65 passed, 1 failed (of 66)
203
+ EXIT=1
204
+ ```
205
+
206
+ Not caused by any held commit. My `autonomy/loki` diff touches only three
207
+ tree-building sites and no pricing or routing code.
208
+
209
+ ### Not fixed here, deliberately
210
+
211
+ Two candidate fixes, and choosing between them is a product call I should not
212
+ make unilaterally:
213
+
214
+ 1. **Enlarge the fixture** so a multi-iteration estimate is exercised. Restores
215
+ the assertion's original intent (verify the architect collapse across a
216
+ real multi-iteration run) and keeps its coverage.
217
+ 2. **Weaken the assertion** to accept `Opus`. Cheaper, and wrong: it would
218
+ stop testing the later-iteration Sonnet attribution entirely.
219
+
220
+ (1) is almost certainly right, but it changes what the test measures, and the
221
+ prior instruction excluded runtime changes without an evidenced deterministic
222
+ requirement. The evidence is now here; the decision is not mine.
223
+
224
+ `docs/LOOP-CANDIDATE-PROPOSAL-v1.md` states the onboard defect "cannot be
225
+ addressed by any of [the six cheaper surfaces]" because it is a bash command
226
+ that never calls a model. That reasoning was right, and the conclusion drawn
227
+ from it was too weak: it is not a model-loop problem, it is a **two-line shell
228
+ bug**, and the correct action was to fix it rather than to route around it.
229
+
230
+ It also claimed the failure needed "a runtime fix to `autonomy/loki`" of
231
+ unknown size. Measured: 28 lines changed across three sites, all mechanical.
232
+
233
+ ## Gate status
234
+
235
+ Still not green, and this triage does not make it so.
236
+
237
+ - `test-onboard-command`: **RESOLVED** (3/9 -> 10/10)
238
+ - `test-model-override`: **1 of 66 still failing, BASELINE** -- a stale test
239
+ coupling, not an estimator defect. Root-caused (fixture no longer produces a
240
+ multi-iteration estimate); two candidate fixes named, neither applied.
241
+ - `bun run typecheck`: **unchanged**, `tsc` still absent
242
+
243
+ Two items remain, not one. Both are pre-existing at `dda8beec`; neither was
244
+ introduced by a held commit.
245
+
246
+ | Item | Needs |
247
+ |---|---|
248
+ | `bun run typecheck` | install the TS toolchain -- the environment fix already named |
249
+ | `test-model-override` | a decision between enlarging the fixture and weakening the assertion |
250
+
251
+ Neither is unblocked by the other, so "install tsc" was never sufficient on its
252
+ own. That was an error in the earlier proposal, which named a single
253
+ gate-closing action while a second real failure sat unclassified behind a
254
+ measurement I had cut short.