loki-mode 9.12.6 → 9.16.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +81 -101
- package/SKILL.md +2 -2
- package/VERSION +1 -1
- package/autonomy/intent.sh +414 -0
- package/autonomy/issue-providers.sh +21 -0
- package/autonomy/lib/agent_readiness.py +202 -0
- package/autonomy/lib/claim_grounding.py +171 -0
- package/autonomy/lib/config-map.sh +10 -6
- package/autonomy/lib/decision_record.py +198 -0
- package/autonomy/lib/failure_memory.py +199 -0
- package/autonomy/lib/outcome_ledger.py +498 -0
- package/autonomy/lib/preedit_snapshot.py +216 -0
- package/autonomy/lib/verdict.py +204 -0
- package/autonomy/loki +358 -14
- package/autonomy/provider-offer.sh +25 -1
- package/autonomy/run.sh +516 -11
- package/autonomy/telemetry.sh +8 -1
- package/completions/_loki +4 -0
- package/completions/loki.bash +2 -1
- package/dashboard/__init__.py +1 -1
- package/dashboard/run.py +13 -2
- package/dashboard/scim.py +221 -0
- package/dashboard/server.py +22 -0
- package/docs/GATE-FAILURE-TRIAGE.md +254 -0
- package/docs/LOOP-CANDIDATE-PROPOSAL-v1.md +167 -0
- package/docs/LOOP-HARNESS-AUDIT.md +53 -0
- package/docs/VERIFICATION-COST.md +103 -0
- package/docs/WANG-PRINCIPLES-PLAN.md +1 -1
- package/loki-ts/dist/loki.js +402 -398
- package/mcp/__init__.py +1 -1
- package/mcp/_sdk_loader.py +25 -0
- package/package.json +1 -1
- package/plugins/loki-mode/.claude-plugin/plugin.json +1 -1
package/autonomy/telemetry.sh
CHANGED
|
@@ -194,7 +194,14 @@ _loki_known_command() {
|
|
|
194
194
|
# a first run, never the user's paths, versions, or hostnames.
|
|
195
195
|
_loki_known_blocker() {
|
|
196
196
|
case "${1:-}" in
|
|
197
|
-
no_provider
|
|
197
|
+
# not_logged_in is distinct from no_provider on purpose. They are the two
|
|
198
|
+
# halves of the same wall and they need opposite fixes: no_provider means
|
|
199
|
+
# "install something", not_logged_in means "authenticate the thing you
|
|
200
|
+
# already installed". Collapsing the second into `other` would make the
|
|
201
|
+
# single most common post-install failure unactionable in the data --
|
|
202
|
+
# and it is the worst-placed one, since without this the user only learns
|
|
203
|
+
# of it after confirming the spend.
|
|
204
|
+
no_provider|not_logged_in|node|python3|jq|git|curl|disk|skill_symlink)
|
|
198
205
|
printf '%s' "$1" ;;
|
|
199
206
|
*)
|
|
200
207
|
printf 'other' ;;
|
package/completions/_loki
CHANGED
|
@@ -171,6 +171,9 @@ function _loki_commands {
|
|
|
171
171
|
'metrics:Session productivity report'
|
|
172
172
|
'share:Share session report as GitHub Gist'
|
|
173
173
|
'proof:Inspect/share proof-of-run artifacts'
|
|
174
|
+
'outcomes:What happened to the work AFTER the receipt (reverted/reworked/survived)'
|
|
175
|
+
'verdict:The five measured trust signals in one readable block'
|
|
176
|
+
'readiness:Can an agent verify its own work in this repo? (measured, not LLM-scored)'
|
|
174
177
|
'preview:Preview the locally-running app'
|
|
175
178
|
'deploy:Deploy the built product (CI/CD-aware)'
|
|
176
179
|
'context:Context window management'
|
|
@@ -190,6 +193,7 @@ function _loki_commands {
|
|
|
190
193
|
'magic:Spec-driven component generation'
|
|
191
194
|
'assets:Export/import shareable team assets'
|
|
192
195
|
'spec:Living spec drift detection'
|
|
196
|
+
'intent:Does the spec still say what was actually wanted'
|
|
193
197
|
'verify:Deterministic PR verification'
|
|
194
198
|
'grill:Interrogate a spec before building'
|
|
195
199
|
'trust:Visible trust trajectory'
|
package/completions/loki.bash
CHANGED
|
@@ -5,7 +5,8 @@ _loki_completion() {
|
|
|
5
5
|
_init_completion || return
|
|
6
6
|
|
|
7
7
|
# Main subcommands (must match autonomy/loki main case statement)
|
|
8
|
-
local main_commands="start quick monitor demo tour welcome init stop pause resume steer status next ship dashboard web serve api sandbox notify import github issue config provider reset memory compound checkpoint council dogfood projects enterprise secrets cockpit secure own handoff doctor watchdog audit metrics syslog onboard share proof receipt explain plan report cost estimate kpis stats test ci watch telemetry agent context ctx code run export review optimize heal modernize migrate cluster worktree wt trigger failover remote deploy docker mcp magic assets analyze compliance crash open otel preview quickstart rc rollback self-update sentrux setup-skill spec state template trust trust-metrics ultracode update verify voice why wiki bench cleanup logs grill docs cp version completions help"
|
|
8
|
+
local main_commands="start quick monitor demo tour welcome init stop pause resume steer status next ship dashboard web serve api sandbox notify import github issue config provider reset memory compound checkpoint council dogfood projects enterprise secrets cockpit secure own handoff doctor watchdog audit metrics syslog onboard share proof receipt outcomes verdict readiness explain plan report cost estimate kpis stats test ci watch telemetry agent context ctx code run export review optimize heal modernize migrate cluster worktree wt trigger failover remote deploy docker mcp magic assets analyze compliance crash open otel preview quickstart rc rollback self-update sentrux setup-skill spec state template trust trust-metrics ultracode update verify voice why wiki bench cleanup logs grill docs cp version completions help"
|
|
9
|
+
local main_commands="start quick monitor demo tour welcome init stop pause resume steer status next ship dashboard web serve api sandbox notify import github issue intent config provider reset memory compound checkpoint council dogfood projects enterprise secrets cockpit secure own handoff doctor watchdog audit metrics syslog onboard share proof receipt explain plan report cost estimate kpis stats test ci watch telemetry agent context ctx code run export review optimize heal modernize migrate cluster worktree wt trigger failover remote deploy docker mcp magic assets analyze compliance crash open otel preview quickstart rc rollback self-update sentrux setup-skill spec state template trust trust-metrics ultracode update verify voice why wiki bench cleanup logs grill docs cp version completions help"
|
|
9
10
|
|
|
10
11
|
# 1. If we are on the first argument (subcommand)
|
|
11
12
|
if [[ $cword -eq 1 ]]; then
|
package/dashboard/__init__.py
CHANGED
package/dashboard/run.py
CHANGED
|
@@ -8,15 +8,26 @@ Usage:
|
|
|
8
8
|
"""
|
|
9
9
|
|
|
10
10
|
import argparse
|
|
11
|
+
import os
|
|
11
12
|
import sys
|
|
12
13
|
|
|
13
14
|
|
|
14
15
|
def main():
|
|
15
16
|
parser = argparse.ArgumentParser(description="Loki Mode Dashboard Server")
|
|
17
|
+
# Loopback by default, matching server.py:run_server and control.py. This
|
|
18
|
+
# launcher defaulted to 0.0.0.0, which published an UNAUTHENTICATED control
|
|
19
|
+
# plane on every interface: dashboard auth is opt-in
|
|
20
|
+
# (LOKI_ENTERPRISE_AUTH, default false in dashboard/auth.py), so with it
|
|
21
|
+
# unset -- the default -- anyone routable to the host could reach endpoints
|
|
22
|
+
# that stop builds and read receipts. A default must be safe on its own;
|
|
23
|
+
# exposing the network is a decision an operator makes explicitly, so
|
|
24
|
+
# --host 0.0.0.0 and LOKI_DASHBOARD_HOST both still work for the container
|
|
25
|
+
# and Helm paths that genuinely need it.
|
|
16
26
|
parser.add_argument(
|
|
17
27
|
"--host",
|
|
18
|
-
default="
|
|
19
|
-
help="Host to bind to (default:
|
|
28
|
+
default=os.environ.get("LOKI_DASHBOARD_HOST", "127.0.0.1"),
|
|
29
|
+
help="Host to bind to (default: $LOKI_DASHBOARD_HOST or 127.0.0.1). "
|
|
30
|
+
"Use 0.0.0.0 only with authentication enabled.",
|
|
20
31
|
)
|
|
21
32
|
parser.add_argument(
|
|
22
33
|
"--port",
|
|
@@ -0,0 +1,221 @@
|
|
|
1
|
+
"""SCIM 2.0 user provisioning: deprovisioning that actually revokes access.
|
|
2
|
+
|
|
3
|
+
WHY THIS EXISTS. Measured against our own code before writing any: OIDC is real
|
|
4
|
+
here (76 references in dashboard/auth.py) and SAML and SCIM are both ZERO. That
|
|
5
|
+
is the whole enterprise identity gap -- not authentication as a category, which
|
|
6
|
+
we have, but PROVISIONING, which we do not.
|
|
7
|
+
|
|
8
|
+
It matters more than its size suggests. With OIDC alone, an employee removed
|
|
9
|
+
from the IdP stops being able to LOG IN, but any token already issued to them
|
|
10
|
+
keeps working until it expires. For a buyer whose security review asks "show me
|
|
11
|
+
that a terminated employee loses access within minutes", that is the answer that
|
|
12
|
+
fails the review. SCIM is the protocol that answers it.
|
|
13
|
+
|
|
14
|
+
WORTH KNOWING ABOUT THE COMPETITION. Factory ships SCIM. Devin's SCIM is
|
|
15
|
+
documented only for Devin DESKTOP -- the page instructs you to configure it at
|
|
16
|
+
windsurf.com/team/settings under the app name "Windsurf" -- and no page documents
|
|
17
|
+
SCIM provisioning for Devin CLOUD sessions, which get SAML/OIDC plus IdP-group
|
|
18
|
+
sync only. So this is table stakes against one competitor and a genuine gap in
|
|
19
|
+
the other.
|
|
20
|
+
|
|
21
|
+
WHAT THIS IS NOT. It is not a second identity system. It maps IdP group
|
|
22
|
+
membership onto the scope hierarchy dashboard/auth.py already enforces
|
|
23
|
+
(* -> control -> write -> read), so there is exactly one authorization model and
|
|
24
|
+
SCIM only decides who holds which scope. A parallel permission model would
|
|
25
|
+
eventually disagree with the real one, and the disagreement would be invisible
|
|
26
|
+
until an audit.
|
|
27
|
+
|
|
28
|
+
DEPROVISIONING IS SOFT BY DEFAULT, AND THAT IS DELIBERATE. `active: false`
|
|
29
|
+
revokes access immediately but PRESERVES the record, because an audit trail that
|
|
30
|
+
loses the identity of whoever performed past actions is not an audit trail. Hard
|
|
31
|
+
delete is available and separate, so destroying history is always an explicit act.
|
|
32
|
+
"""
|
|
33
|
+
|
|
34
|
+
from __future__ import annotations
|
|
35
|
+
|
|
36
|
+
import json
|
|
37
|
+
import os
|
|
38
|
+
from datetime import datetime, timezone
|
|
39
|
+
from typing import Optional
|
|
40
|
+
|
|
41
|
+
SCIM_SCHEMA_USER = "urn:ietf:params:scim:schemas:core:2.0:User"
|
|
42
|
+
SCIM_SCHEMA_LIST = "urn:ietf:params:scim:api:messages:2.0:ListResponse"
|
|
43
|
+
SCIM_SCHEMA_ERROR = "urn:ietf:params:scim:api:messages:2.0:Error"
|
|
44
|
+
|
|
45
|
+
# IdP group -> our scope. Mirrors the hierarchy in dashboard/auth.py rather than
|
|
46
|
+
# inventing a second one: * grants everything, control grants write and read.
|
|
47
|
+
# An unrecognised group maps to "read", never to a higher scope -- an unknown
|
|
48
|
+
# group is not evidence of privilege, and defaulting upward is how a
|
|
49
|
+
# misconfigured IdP silently grants admin.
|
|
50
|
+
GROUP_SCOPE_MAP = {
|
|
51
|
+
"loki-admins": "*",
|
|
52
|
+
"loki-owners": "*",
|
|
53
|
+
"loki-operators": "control",
|
|
54
|
+
"loki-writers": "write",
|
|
55
|
+
"loki-readers": "read",
|
|
56
|
+
}
|
|
57
|
+
DEFAULT_SCOPE = "read"
|
|
58
|
+
|
|
59
|
+
|
|
60
|
+
def _now():
|
|
61
|
+
return datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
|
|
62
|
+
|
|
63
|
+
|
|
64
|
+
def _store_path(loki_dir):
|
|
65
|
+
return os.path.join(loki_dir, "scim", "users.json")
|
|
66
|
+
|
|
67
|
+
|
|
68
|
+
def _load(loki_dir):
|
|
69
|
+
p = _store_path(loki_dir)
|
|
70
|
+
if not os.path.isfile(p):
|
|
71
|
+
return {}
|
|
72
|
+
try:
|
|
73
|
+
with open(p, "r", encoding="utf-8") as fh:
|
|
74
|
+
return json.load(fh)
|
|
75
|
+
except (OSError, ValueError):
|
|
76
|
+
# A corrupt store must not silently become an empty one: returning {}
|
|
77
|
+
# here would read as "no users provisioned" and could re-grant access to
|
|
78
|
+
# someone who was deprovisioned. Signalled to the caller instead.
|
|
79
|
+
return None
|
|
80
|
+
|
|
81
|
+
|
|
82
|
+
def _save(loki_dir, data):
|
|
83
|
+
p = _store_path(loki_dir)
|
|
84
|
+
os.makedirs(os.path.dirname(p), exist_ok=True)
|
|
85
|
+
tmp = p + ".tmp"
|
|
86
|
+
with open(tmp, "w", encoding="utf-8") as fh:
|
|
87
|
+
json.dump(data, fh, indent=2, sort_keys=True)
|
|
88
|
+
fh.write("\n")
|
|
89
|
+
# Atomic replace: a half-written user store is an access-control decision
|
|
90
|
+
# made from corrupt data.
|
|
91
|
+
os.replace(tmp, p)
|
|
92
|
+
|
|
93
|
+
|
|
94
|
+
def scope_for_groups(groups):
|
|
95
|
+
"""Highest scope any of the user's groups grants.
|
|
96
|
+
|
|
97
|
+
Highest-wins, because a user in both loki-readers and loki-admins is an
|
|
98
|
+
admin. Unknown groups contribute DEFAULT_SCOPE and never more.
|
|
99
|
+
"""
|
|
100
|
+
rank = {"read": 1, "write": 2, "control": 3, "*": 4}
|
|
101
|
+
best = DEFAULT_SCOPE
|
|
102
|
+
for g in groups or []:
|
|
103
|
+
s = GROUP_SCOPE_MAP.get(str(g).strip().lower(), DEFAULT_SCOPE)
|
|
104
|
+
if rank.get(s, 1) > rank.get(best, 1):
|
|
105
|
+
best = s
|
|
106
|
+
return best
|
|
107
|
+
|
|
108
|
+
|
|
109
|
+
def provision(loki_dir, user_name, external_id=None, groups=None, active=True):
|
|
110
|
+
"""Create or update a user. Returns a SCIM User resource."""
|
|
111
|
+
if not user_name:
|
|
112
|
+
return {"schemas": [SCIM_SCHEMA_ERROR], "status": "400",
|
|
113
|
+
"detail": "userName is required"}
|
|
114
|
+
|
|
115
|
+
data = _load(loki_dir)
|
|
116
|
+
if data is None:
|
|
117
|
+
return {"schemas": [SCIM_SCHEMA_ERROR], "status": "500",
|
|
118
|
+
"detail": "user store is unreadable; refusing to make an access "
|
|
119
|
+
"decision from corrupt data"}
|
|
120
|
+
|
|
121
|
+
uid = external_id or user_name
|
|
122
|
+
existing = data.get(uid, {})
|
|
123
|
+
rec = {
|
|
124
|
+
"id": uid,
|
|
125
|
+
"userName": user_name,
|
|
126
|
+
"externalId": external_id or "",
|
|
127
|
+
"groups": list(groups or []),
|
|
128
|
+
"scope": scope_for_groups(groups),
|
|
129
|
+
"active": bool(active),
|
|
130
|
+
"created": existing.get("created") or _now(),
|
|
131
|
+
"lastModified": _now(),
|
|
132
|
+
}
|
|
133
|
+
data[uid] = rec
|
|
134
|
+
_save(loki_dir, data)
|
|
135
|
+
return to_scim_user(rec)
|
|
136
|
+
|
|
137
|
+
|
|
138
|
+
def deprovision(loki_dir, uid, hard=False):
|
|
139
|
+
"""Revoke access. Soft by default so the audit trail survives.
|
|
140
|
+
|
|
141
|
+
An audit log that cannot say WHO performed an action is not an audit log, so
|
|
142
|
+
a departed user's record is retained by default and merely marked inactive.
|
|
143
|
+
Hard delete exists and is separate, so destroying history is always explicit.
|
|
144
|
+
"""
|
|
145
|
+
data = _load(loki_dir)
|
|
146
|
+
if data is None:
|
|
147
|
+
return {"schemas": [SCIM_SCHEMA_ERROR], "status": "500",
|
|
148
|
+
"detail": "user store is unreadable"}
|
|
149
|
+
if uid not in data:
|
|
150
|
+
return {"schemas": [SCIM_SCHEMA_ERROR], "status": "404",
|
|
151
|
+
"detail": f"no such user: {uid}"}
|
|
152
|
+
|
|
153
|
+
if hard:
|
|
154
|
+
data.pop(uid)
|
|
155
|
+
_save(loki_dir, data)
|
|
156
|
+
return {"status": "204", "deleted": uid, "audit_trail_preserved": False}
|
|
157
|
+
|
|
158
|
+
data[uid]["active"] = False
|
|
159
|
+
data[uid]["lastModified"] = _now()
|
|
160
|
+
_save(loki_dir, data)
|
|
161
|
+
return {"status": "200", "deactivated": uid, "audit_trail_preserved": True}
|
|
162
|
+
|
|
163
|
+
|
|
164
|
+
def is_active(loki_dir, uid):
|
|
165
|
+
"""Authorization answer for one user. Fails CLOSED on any doubt.
|
|
166
|
+
|
|
167
|
+
An unreadable store, an unknown user, and an inactive user all return False.
|
|
168
|
+
A provisioning system that fails open is worse than none, because it creates
|
|
169
|
+
the belief that deprovisioning works.
|
|
170
|
+
"""
|
|
171
|
+
data = _load(loki_dir)
|
|
172
|
+
if data is None:
|
|
173
|
+
return False
|
|
174
|
+
rec = data.get(uid)
|
|
175
|
+
if not rec:
|
|
176
|
+
return False
|
|
177
|
+
return bool(rec.get("active"))
|
|
178
|
+
|
|
179
|
+
|
|
180
|
+
def effective_scope(loki_dir, uid):
|
|
181
|
+
"""The scope a user actually holds, or None when they hold none."""
|
|
182
|
+
data = _load(loki_dir)
|
|
183
|
+
if data is None:
|
|
184
|
+
return None
|
|
185
|
+
rec = data.get(uid)
|
|
186
|
+
if not rec or not rec.get("active"):
|
|
187
|
+
return None
|
|
188
|
+
return rec.get("scope") or DEFAULT_SCOPE
|
|
189
|
+
|
|
190
|
+
|
|
191
|
+
def to_scim_user(rec):
|
|
192
|
+
return {
|
|
193
|
+
"schemas": [SCIM_SCHEMA_USER],
|
|
194
|
+
"id": rec["id"],
|
|
195
|
+
"userName": rec["userName"],
|
|
196
|
+
"externalId": rec.get("externalId", ""),
|
|
197
|
+
"active": rec.get("active", False),
|
|
198
|
+
"groups": [{"value": g} for g in rec.get("groups", [])],
|
|
199
|
+
"meta": {
|
|
200
|
+
"resourceType": "User",
|
|
201
|
+
"created": rec.get("created"),
|
|
202
|
+
"lastModified": rec.get("lastModified"),
|
|
203
|
+
},
|
|
204
|
+
# Not part of the SCIM spec, surfaced because an operator reading this
|
|
205
|
+
# response needs to see what the group mapping actually granted rather
|
|
206
|
+
# than deriving it.
|
|
207
|
+
"lokiScope": rec.get("scope"),
|
|
208
|
+
}
|
|
209
|
+
|
|
210
|
+
|
|
211
|
+
def list_users(loki_dir, active_only=False):
|
|
212
|
+
data = _load(loki_dir)
|
|
213
|
+
if data is None:
|
|
214
|
+
return {"schemas": [SCIM_SCHEMA_ERROR], "status": "500",
|
|
215
|
+
"detail": "user store is unreadable"}
|
|
216
|
+
recs = [r for r in data.values() if (not active_only or r.get("active"))]
|
|
217
|
+
return {
|
|
218
|
+
"schemas": [SCIM_SCHEMA_LIST],
|
|
219
|
+
"totalResults": len(recs),
|
|
220
|
+
"Resources": [to_scim_user(r) for r in sorted(recs, key=lambda r: r["id"])],
|
|
221
|
+
}
|
package/dashboard/server.py
CHANGED
|
@@ -1224,6 +1224,28 @@ class WebSocketBoundaryMiddleware:
|
|
|
1224
1224
|
# accept-then-drop, which reads to a client as a flaky network.
|
|
1225
1225
|
await send({"type": "websocket.close", "code": 1008})
|
|
1226
1226
|
return
|
|
1227
|
+
# Cross-Site WebSocket Hijacking. The address check above proves the
|
|
1228
|
+
# PEER is local; it says nothing about who told that peer to connect.
|
|
1229
|
+
# A browser on any origin can open ws://127.0.0.1:57374/ws -- the
|
|
1230
|
+
# request comes FROM the loopback interface and passes the check
|
|
1231
|
+
# above, and CORS does not apply to WebSocket upgrades, so no
|
|
1232
|
+
# existing control stops it. /ws/collab is writable, so this is not
|
|
1233
|
+
# read-only exposure.
|
|
1234
|
+
#
|
|
1235
|
+
# Same-origin browsers send Origin; non-browser clients (CLI, tests,
|
|
1236
|
+
# health probes) send none. An ABSENT Origin is therefore allowed --
|
|
1237
|
+
# it cannot be forged by a page, since browsers always attach it --
|
|
1238
|
+
# while a PRESENT one must be in the allowlist the HTTP path already
|
|
1239
|
+
# uses. That keeps every scripted client working and closes the
|
|
1240
|
+
# browser-driven path, without a second source of truth for origins.
|
|
1241
|
+
origin = None
|
|
1242
|
+
for raw_name, raw_value in scope.get("headers", []):
|
|
1243
|
+
if raw_name == b"origin":
|
|
1244
|
+
origin = raw_value.decode("latin-1").strip()
|
|
1245
|
+
break
|
|
1246
|
+
if origin and "*" not in _cors_origins and origin not in _cors_origins:
|
|
1247
|
+
await send({"type": "websocket.close", "code": 1008})
|
|
1248
|
+
return
|
|
1227
1249
|
await self.app(scope, receive, send)
|
|
1228
1250
|
|
|
1229
1251
|
|
|
@@ -0,0 +1,254 @@
|
|
|
1
|
+
# Gate failure triage: exact-SHA classification
|
|
2
|
+
|
|
3
|
+
Ordered by the steer: reproduce locally against exact HEAD, classify each
|
|
4
|
+
failure as environment / baseline / candidate regression, evidence every claim
|
|
5
|
+
with a command and its output. Nothing here was pushed.
|
|
6
|
+
|
|
7
|
+
| Field | Value |
|
|
8
|
+
|---|---|
|
|
9
|
+
| HEAD at triage | `d230a3b2` (the steer named `09138e26`; that SHA is not in this worktree) |
|
|
10
|
+
| baseline compared | `dda8beec` = origin/main |
|
|
11
|
+
| held commits | `22199024`, `94315f35`, `d230a3b2`, plus `94f7639a` from this triage |
|
|
12
|
+
| pushed | **nothing** |
|
|
13
|
+
|
|
14
|
+
## Classification
|
|
15
|
+
|
|
16
|
+
| Failure | Class | Evidence |
|
|
17
|
+
|---|---|---|
|
|
18
|
+
| `test-onboard-command` (6 of 9) | **BASELINE**, now FIXED | `autonomy/loki` byte-identical to origin/main; `git diff --name-only dda8beec..HEAD` returns zero matches for that path |
|
|
19
|
+
| `test-model-override` | **BASELINE**, 1 of 66, still open | identical failure at `dda8beec` and at HEAD: `Results: 65 passed, 1 failed (of 66)`, `EXIT=1` |
|
|
20
|
+
| `bun run typecheck` | **ENVIRONMENT** | `tsc` not installed locally; unchanged |
|
|
21
|
+
|
|
22
|
+
## The onboard defect
|
|
23
|
+
|
|
24
|
+
`loki onboard --stdout` exited **141** and wrote **0 bytes** on this repo,
|
|
25
|
+
while passing on small fixtures.
|
|
26
|
+
|
|
27
|
+
```
|
|
28
|
+
EXIT=141
|
|
29
|
+
STDOUT bytes: 0
|
|
30
|
+
STDERR bytes: 242
|
|
31
|
+
```
|
|
32
|
+
|
|
33
|
+
141 = 128+13 = SIGPIPE. `cmd_onboard`'s find fallback ends
|
|
34
|
+
`| sed | sort | head -200`. Under the file's `set -euo pipefail` (line 22),
|
|
35
|
+
`head` exits after 200 lines; `sort` -- which must consume ALL input before it
|
|
36
|
+
emits anything -- then takes SIGPIPE, the pipeline returns 141, and `-e` aborts
|
|
37
|
+
before a byte is written. The sibling `git ls-files` branch two lines above was
|
|
38
|
+
already guarded with `|| true`. The find branch never was.
|
|
39
|
+
|
|
40
|
+
`sort`, not `find`, is the process that dies. That distinction sets the test
|
|
41
|
+
size: see below.
|
|
42
|
+
|
|
43
|
+
### Why no existing test caught it
|
|
44
|
+
|
|
45
|
+
**1. Every fixture was too small.** The residual output has to exceed the 64KB
|
|
46
|
+
pipe buffer before the signal lands. Measured:
|
|
47
|
+
|
|
48
|
+
| Fixture | Result vs unfixed code |
|
|
49
|
+
|---|---|
|
|
50
|
+
| 250 files | **passes** -- ~50 lines left after the cut, fits the buffer |
|
|
51
|
+
| 3000 files | **exit 141**, deterministic |
|
|
52
|
+
|
|
53
|
+
A 250-file regression test would have been worthless. I wrote one first,
|
|
54
|
+
confirmed it passed against the pre-fix binary, and resized it.
|
|
55
|
+
|
|
56
|
+
**2. This worktree never reached the guarded branch.** The check was
|
|
57
|
+
`[ -d "$target_path/.git" ]`, and in a git **worktree** `.git` is a pointer
|
|
58
|
+
**file**, not a directory:
|
|
59
|
+
|
|
60
|
+
```
|
|
61
|
+
-rw-r--r-- 1 lokesh staff 83 Jul 31 19:25 .git
|
|
62
|
+
gitdir: /Users/lokesh/git/lokimode-anthropic/.git/worktrees/pre-push-scoped-pytest
|
|
63
|
+
```
|
|
64
|
+
|
|
65
|
+
So every worktree silently fell through to the find path. Proven by trace:
|
|
66
|
+
|
|
67
|
+
```
|
|
68
|
+
PRE-FIX ++ find ... -maxdepth 4
|
|
69
|
+
FIXED ++ git ls-files
|
|
70
|
+
```
|
|
71
|
+
|
|
72
|
+
### The fix
|
|
73
|
+
|
|
74
|
+
`-e` instead of `-d`, and `|| true` matching the sibling branch. Two lines.
|
|
75
|
+
|
|
76
|
+
### Two sibling sites, quieter symptom
|
|
77
|
+
|
|
78
|
+
`cmd_explain` and `_docs_scan_project` carry the same pipeline at `head -500`.
|
|
79
|
+
Fixed alongside -- patching only the path the failure named would leave the
|
|
80
|
+
siblings broken.
|
|
81
|
+
|
|
82
|
+
They fail *differently*, which is why nothing ever caught them: both assign via
|
|
83
|
+
`local x=$(...)`, and `local` resets `$?`, swallowing the 141. Demonstrated:
|
|
84
|
+
|
|
85
|
+
```
|
|
86
|
+
$ f() { local x=$(false | head -1); echo "rc=$?"; }
|
|
87
|
+
rc=0
|
|
88
|
+
```
|
|
89
|
+
|
|
90
|
+
So they **silently truncate** their file tree instead of aborting. Same root
|
|
91
|
+
cause, no visible symptom.
|
|
92
|
+
|
|
93
|
+
### Sweep
|
|
94
|
+
|
|
95
|
+
Three unguarded `sort | head -N` sites existed; zero remain. The sweep pattern
|
|
96
|
+
is not vacuous -- it matches 3 in the pre-fix file and 0 now.
|
|
97
|
+
|
|
98
|
+
The other two `-d .../.git` checks in the file (`loki:11113`, `loki:13473`) are
|
|
99
|
+
CORRECT as `-d`: one detects a clone (a worktree is not one), the other guards
|
|
100
|
+
`git init` on a fresh demo dir. Left alone.
|
|
101
|
+
|
|
102
|
+
## Verification
|
|
103
|
+
|
|
104
|
+
| Check | Before | After |
|
|
105
|
+
|---|---|---|
|
|
106
|
+
| `test-onboard-command.sh` | 3/9 | **10/10** |
|
|
107
|
+
| new Test 10 vs pre-fix binary | **FAIL** (exit 141) | PASS |
|
|
108
|
+
| `test-onboard-json-injection-wave10.sh` | 2/2 | 2/2 |
|
|
109
|
+
| `test-contradiction-detection.sh` | 19/19 | 19/19 |
|
|
110
|
+
| `bash -n autonomy/loki` | OK | OK |
|
|
111
|
+
|
|
112
|
+
Test 10 was mutation-tested against `dda8beec`: it fails with the exact
|
|
113
|
+
diagnostic `exit 141 (SIGPIPE)` on the old code and passes on the new. It also
|
|
114
|
+
carries a vacuity guard rejecting exit 0 with under 100 bytes of output -- the
|
|
115
|
+
precise shape of the bug, since the abort produced exit 141 *and* silence.
|
|
116
|
+
|
|
117
|
+
## Correction: I called test-model-override a non-failure before it finished
|
|
118
|
+
|
|
119
|
+
An earlier revision of THIS FILE classified `test-model-override` as "NOT A
|
|
120
|
+
FAILURE -- slow suite, mis-measured". That was wrong, and it was wrong in the
|
|
121
|
+
worst available way: I wrote the classification while the run was still
|
|
122
|
+
executing, from a partial log that showed 50 PASS and no failures yet.
|
|
123
|
+
|
|
124
|
+
The completed run:
|
|
125
|
+
|
|
126
|
+
```
|
|
127
|
+
FAIL: architect no-cap mismatch: estimator='Opus' runner-dispatch='opus'
|
|
128
|
+
runner-tier='fable' (expected Opus,Sonnet / opus / fable)
|
|
129
|
+
Results: 65 passed, 1 failed (of 66)
|
|
130
|
+
EXIT=1
|
|
131
|
+
```
|
|
132
|
+
|
|
133
|
+
The slowness was real -- a 224-cell parity matrix, each cell spawning a
|
|
134
|
+
`python3` that imports the FastAPI dashboard at ~0.46s -- and it was NOT the
|
|
135
|
+
explanation for the failure. Both things were true and I reported only the
|
|
136
|
+
convenient one.
|
|
137
|
+
|
|
138
|
+
The rule this violates is one already written down in this repo: an absent
|
|
139
|
+
measurement is not a measurement. A log with no FAIL line yet is not a log with
|
|
140
|
+
no failures; it is an unfinished log. I should have blocked on the EXIT marker
|
|
141
|
+
before classifying, exactly as I did for the onboard suite.
|
|
142
|
+
|
|
143
|
+
### The actual failure
|
|
144
|
+
|
|
145
|
+
`tests/test-model-override.sh:827`. A three-way coherence assertion; two of the
|
|
146
|
+
three legs are correct:
|
|
147
|
+
|
|
148
|
+
| Leg | Expected | Actual |
|
|
149
|
+
|---|---|---|
|
|
150
|
+
| runner dispatch | `opus` | `opus` -- correct |
|
|
151
|
+
| runner tier (pre-collapse) | `fable` | `fable` -- correct |
|
|
152
|
+
| estimator quote | `Opus,Sonnet` | `Opus` -- **mismatch** |
|
|
153
|
+
|
|
154
|
+
So the runtime routing is right and only the cost QUOTE disagrees: the
|
|
155
|
+
estimator names one model where the run actually uses two. Per the test's own
|
|
156
|
+
comment (line 800), this is the known "estimator needs the sonnet5-default
|
|
157
|
+
update" case -- iter-1 collapses fable to opus, later iterations run the
|
|
158
|
+
development tier which defaults to sonnet since v7.104.0, so an honest quote
|
|
159
|
+
must name both.
|
|
160
|
+
|
|
161
|
+
It under-quotes cost. It does not mis-route a model.
|
|
162
|
+
|
|
163
|
+
### Root cause: the fixture stopped producing enough iterations
|
|
164
|
+
|
|
165
|
+
The estimator is CORRECT. The expectation is only reachable when the estimate
|
|
166
|
+
spans more than one iteration, and the suite's fixture no longer does.
|
|
167
|
+
|
|
168
|
+
`autonomy/loki:18342` prices iteration 0 as Opus (the fable architect pass
|
|
169
|
+
collapsing to opus), and every LATER iteration through
|
|
170
|
+
`_priced_model_for(_dispatched_model)`, which defaults to Sonnet since
|
|
171
|
+
v7.104.0. So `Opus,Sonnet` requires **iterations >= 2**.
|
|
172
|
+
|
|
173
|
+
Measured on the suite's own fixture (`# PRD\nBuild a small todo API with one
|
|
174
|
+
endpoint.`, byte-identical to v7.104.0):
|
|
175
|
+
|
|
176
|
+
| Binary | tier | estimated iterations | nonzero models |
|
|
177
|
+
|---|---|---|---|
|
|
178
|
+
| `766219ac` (v7.104.0, where this was written and passed 66/0) | simple | **4** | `Opus,Sonnet` |
|
|
179
|
+
| HEAD | simple | **1** | `Opus` |
|
|
180
|
+
|
|
181
|
+
The complexity TIER is unchanged (`simple` in both). Only the iteration count
|
|
182
|
+
for that tier fell, 4 -> 1, and with a single iteration the loop never reaches
|
|
183
|
+
the branch where Sonnet appears.
|
|
184
|
+
|
|
185
|
+
Confirmed causal by holding the binary fixed and enlarging the input: a 24-
|
|
186
|
+
feature PRD at HEAD estimates 4 iterations and returns exactly `Opus,Sonnet`
|
|
187
|
+
(`{"Fable":0,"Opus":1,"Sonnet":3,"Haiku":0}`). Same code, more iterations,
|
|
188
|
+
expected answer.
|
|
189
|
+
|
|
190
|
+
So the assertion is a **stale coupling**: it encodes "a simple PRD takes
|
|
191
|
+
several iterations", which stopped being true. The v7.104.0 commit message
|
|
192
|
+
claims "locked by tests/test-model-override.sh (66/0)" -- that lock silently
|
|
193
|
+
came undone when the iteration estimate for simple PRDs changed.
|
|
194
|
+
|
|
195
|
+
### Classification: BASELINE
|
|
196
|
+
|
|
197
|
+
Run against `dda8beec`'s `autonomy/loki` (my onboard fix reverted), the failure
|
|
198
|
+
is **identical**:
|
|
199
|
+
|
|
200
|
+
```
|
|
201
|
+
FAIL: architect no-cap mismatch: estimator='Opus' ...
|
|
202
|
+
Results: 65 passed, 1 failed (of 66)
|
|
203
|
+
EXIT=1
|
|
204
|
+
```
|
|
205
|
+
|
|
206
|
+
Not caused by any held commit. My `autonomy/loki` diff touches only three
|
|
207
|
+
tree-building sites and no pricing or routing code.
|
|
208
|
+
|
|
209
|
+
### Not fixed here, deliberately
|
|
210
|
+
|
|
211
|
+
Two candidate fixes, and choosing between them is a product call I should not
|
|
212
|
+
make unilaterally:
|
|
213
|
+
|
|
214
|
+
1. **Enlarge the fixture** so a multi-iteration estimate is exercised. Restores
|
|
215
|
+
the assertion's original intent (verify the architect collapse across a
|
|
216
|
+
real multi-iteration run) and keeps its coverage.
|
|
217
|
+
2. **Weaken the assertion** to accept `Opus`. Cheaper, and wrong: it would
|
|
218
|
+
stop testing the later-iteration Sonnet attribution entirely.
|
|
219
|
+
|
|
220
|
+
(1) is almost certainly right, but it changes what the test measures, and the
|
|
221
|
+
prior instruction excluded runtime changes without an evidenced deterministic
|
|
222
|
+
requirement. The evidence is now here; the decision is not mine.
|
|
223
|
+
|
|
224
|
+
`docs/LOOP-CANDIDATE-PROPOSAL-v1.md` states the onboard defect "cannot be
|
|
225
|
+
addressed by any of [the six cheaper surfaces]" because it is a bash command
|
|
226
|
+
that never calls a model. That reasoning was right, and the conclusion drawn
|
|
227
|
+
from it was too weak: it is not a model-loop problem, it is a **two-line shell
|
|
228
|
+
bug**, and the correct action was to fix it rather than to route around it.
|
|
229
|
+
|
|
230
|
+
It also claimed the failure needed "a runtime fix to `autonomy/loki`" of
|
|
231
|
+
unknown size. Measured: 28 lines changed across three sites, all mechanical.
|
|
232
|
+
|
|
233
|
+
## Gate status
|
|
234
|
+
|
|
235
|
+
Still not green, and this triage does not make it so.
|
|
236
|
+
|
|
237
|
+
- `test-onboard-command`: **RESOLVED** (3/9 -> 10/10)
|
|
238
|
+
- `test-model-override`: **1 of 66 still failing, BASELINE** -- a stale test
|
|
239
|
+
coupling, not an estimator defect. Root-caused (fixture no longer produces a
|
|
240
|
+
multi-iteration estimate); two candidate fixes named, neither applied.
|
|
241
|
+
- `bun run typecheck`: **unchanged**, `tsc` still absent
|
|
242
|
+
|
|
243
|
+
Two items remain, not one. Both are pre-existing at `dda8beec`; neither was
|
|
244
|
+
introduced by a held commit.
|
|
245
|
+
|
|
246
|
+
| Item | Needs |
|
|
247
|
+
|---|---|
|
|
248
|
+
| `bun run typecheck` | install the TS toolchain -- the environment fix already named |
|
|
249
|
+
| `test-model-override` | a decision between enlarging the fixture and weakening the assertion |
|
|
250
|
+
|
|
251
|
+
Neither is unblocked by the other, so "install tsc" was never sufficient on its
|
|
252
|
+
own. That was an error in the earlier proposal, which named a single
|
|
253
|
+
gate-closing action while a second real failure sat unclassified behind a
|
|
254
|
+
measurement I had cut short.
|