loki-mode 9.8.0 → 9.12.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (79) hide show
  1. package/README.md +19 -14
  2. package/SKILL.md +3 -2
  3. package/VERSION +1 -1
  4. package/autonomy/loki +122 -1
  5. package/autonomy/run.sh +49 -2
  6. package/dashboard/__init__.py +1 -1
  7. package/dashboard/api_evidence.py +411 -0
  8. package/dashboard/api_operator.py +283 -0
  9. package/dashboard/api_phases.py +262 -0
  10. package/dashboard/api_releases.py +242 -0
  11. package/dashboard/api_runs.py +477 -0
  12. package/dashboard/api_tests.py +444 -0
  13. package/dashboard/api_v2.py +47 -1
  14. package/dashboard/server.py +54 -0
  15. package/dashboard/static/index.html +246 -135
  16. package/docs/ARCHITECTURE-OVERVIEW.md +5 -3
  17. package/docs/CAPABILITY-BACKLOG.md +53 -0
  18. package/docs/COMPARISON.md +2 -2
  19. package/docs/COMPETITIVE-ANALYSIS.md +1 -1
  20. package/docs/COMPETITIVE-SCORECARD.md +422 -0
  21. package/docs/DASHBOARD-9.12-EVIDENCE.md +97 -0
  22. package/docs/DASHBOARD-ARCHITECTURE.md +423 -0
  23. package/docs/DEMOS.md +21 -23
  24. package/docs/HANDOFF-2026-08-03.md +439 -0
  25. package/docs/INSTALLATION.md +17 -10
  26. package/docs/OUTCOME-FRONTIER.md +536 -0
  27. package/docs/PROMPT-ABLATION-RESULT.md +97 -0
  28. package/docs/TOOLS.md +800 -0
  29. package/docs/alternative-installations.md +2 -3
  30. package/docs/audit-logging.md +44 -35
  31. package/docs/authentication.md +13 -2
  32. package/docs/authorization.md +87 -81
  33. package/docs/git-workflow.md +6 -3
  34. package/docs/metrics.md +15 -16
  35. package/docs/network-security.md +16 -13
  36. package/docs/openclaw-integration.md +36 -556
  37. package/docs/show-hn-post.md +2 -2
  38. package/docs/siem-integration.md +39 -36
  39. package/loki-ts/dist/loki.js +18 -18
  40. package/mcp/__init__.py +1 -1
  41. package/package.json +2 -2
  42. package/plugins/loki-mode/.claude-plugin/plugin.json +1 -1
  43. package/references/confidence-routing.md +18 -1
  44. package/references/invariant-checks.md +13 -8
  45. package/references/magic-rarv-integration.md +0 -1
  46. package/references/multi-provider.md +27 -5
  47. package/skills/healing.md +4 -2
  48. package/tools/audit-docs.py +488 -0
  49. package/tools/baseline-pin.py +19 -1
  50. package/tools/calibration-audit.py +523 -0
  51. package/tools/ci-gate.py +19 -1
  52. package/tools/cost-forecast.py +344 -0
  53. package/tools/cost-guard.py +19 -1
  54. package/tools/cost-history.py +19 -1
  55. package/tools/cost-per-outcome.py +394 -0
  56. package/tools/estimate-run.py +19 -1
  57. package/tools/evidence-freshness.py +307 -0
  58. package/tools/gate-init.py +19 -1
  59. package/tools/gate-report.py +19 -1
  60. package/tools/gate-simulate.py +570 -0
  61. package/tools/gate-trend.py +354 -0
  62. package/tools/model-advisor.py +52 -1
  63. package/tools/policy-load.py +19 -1
  64. package/tools/prompt-cost.py +363 -0
  65. package/tools/prompt-diff.py +448 -0
  66. package/tools/prompt-lint.py +448 -0
  67. package/tools/receipt-bundle.py +72 -2
  68. package/tools/receipt-diff.py +19 -1
  69. package/tools/receipt-find.py +19 -1
  70. package/tools/receipt-stats.py +380 -0
  71. package/tools/receipt-timeline.py +478 -0
  72. package/tools/receipt-verify-batch.py +291 -0
  73. package/tools/run-replay.py +19 -1
  74. package/tools/signing-status.py +19 -1
  75. package/tools/token-guard.py +19 -1
  76. package/tools/token-tax.py +375 -0
  77. package/tools/tool-index.py +19 -1
  78. package/tools/verification-tax.py +277 -0
  79. package/tools/verify-chain.py +361 -0
@@ -0,0 +1,423 @@
1
+ # Dashboard Architecture: Real-Data, Near-Realtime
2
+
3
+ Design only. No implementation in this document.
4
+
5
+ Status: proposal. Measured against the worktree at
6
+ `/Users/lokesh/git/lokimode-anthropic/.claude/worktrees/pre-push-scoped-pytest`
7
+ on 2026-08-03.
8
+
9
+ ## 0. Executive finding
10
+
11
+ The dashboard's problem is not styling and not the UI framework. It is that
12
+ several required domains have **no read path from the engine to the browser**.
13
+
14
+ But the brief's framing ("six domains have ZERO API surface") is wrong in a way
15
+ that makes the work cheaper and changes its ranking. The precise situation is
16
+ three different failure modes, and they need three different fixes:
17
+
18
+ | Failure mode | Domains | Fix shape | Cost |
19
+ |---|---|---|---|
20
+ | **A. Route exists, store is never written** | runs | Wire an engine-side writer | Small |
21
+ | **B. Writer exists in code, no route reads it** | receipts, tests, artifacts, prompts (partial), releases | Add a read route over the file that writer produces | Small each |
22
+ | **C. No source of truth at all** | prompts (main-loop) | Do not build. Mark NOT AVAILABLE | Zero |
23
+
24
+ **Precision on mode B.** What is verified is that a *writer exists in source* at
25
+ a cited `file:line`. It is **not** verified that an instance exists on disk: this
26
+ worktree has never completed a run, `.loki/proofs/` and `.loki/quality/` do not
27
+ exist here, and `ls -R .loki/proofs` returns empty. Claiming "the data is
28
+ already there" would be exactly the fabrication this document forbids (see R8).
29
+ The design conclusion is unchanged, because a read route is written against the
30
+ writer's contract, not against a sample file. But the two facts are tracked
31
+ separately throughout, and every mode-B row below says which one it is.
32
+
33
+ None of these is a UI rewrite. The ranked work is **writer-side and
34
+ route-side**. Section 4 concludes that the existing 46-component UI is kept
35
+ almost entirely, and that a rewrite is not supportable on the evidence.
36
+
37
+ ### Measurement honesty
38
+
39
+ Every number below was produced by a command recorded next to it. Where the
40
+ brief's number and my measurement disagree, both are shown.
41
+
42
+ One correction against myself: my first pass reported `0` aria attributes in the
43
+ built artifact and I nearly wrote that up as a shipped accessibility regression.
44
+ It was a measurement error. `dashboard/static/index.html` has lines long enough
45
+ (577 chars plus minified runs) that `grep` treated the file as binary and
46
+ suppressed output. With `grep -a` the brief's numbers are confirmed. This is the
47
+ `grep-absence false-green` failure mode: an empty result was an absent
48
+ measurement, not evidence of absence. Any count in this document taken over a
49
+ built artifact uses `grep -a`.
50
+
51
+ ## 1. DATA CONTRACT
52
+
53
+ ### 1.1 How route counts were measured
54
+
55
+ ```
56
+ grep -oE '^@app\.(get|post|put|delete|patch|websocket)\("([^"]+)"' \
57
+ dashboard/server.py | sed 's/^@app\.\([a-z]*\)("/\1 /' | sort -u # 156
58
+ grep -cE '^@app\.(get|post|put|delete|patch|websocket)' dashboard/server.py # 165
59
+ ```
60
+
61
+ - 165 route decorators in `dashboard/server.py`
62
+ - 156 unique (method, path) pairs
63
+ - 140 unique paths
64
+ - Plus 24 routes in `dashboard/api_v2.py` under prefix `/api/v2`
65
+ (`dashboard/api_v2.py:39`), mounted at `dashboard/server.py:1010-1011`
66
+
67
+ The brief's "138 routes" is close to the 140 unique-path figure but is not
68
+ reproducible from any single command I ran. **The api_v2 router was not counted
69
+ in the brief at all**, which is how the runs domain got misclassified.
70
+
71
+ `dashboard/server.py` is 12,466 lines.
72
+
73
+ ### 1.2 The 15 domains
74
+
75
+ Freshness classes used below:
76
+
77
+ - **PUSH** - broadcast by `_push_loki_state_loop` (`dashboard/server.py:596`),
78
+ 2s while a run is active, 30s idle, 5s with no clients connected
79
+ - **MTIME** - file modification time is available and should be returned
80
+ - **ON-READ** - computed per request, freshness equals request time
81
+ - **NONE** - no freshness signal exists
82
+
83
+ | # | Domain | API today | Source of truth on disk | Freshness | Verdict |
84
+ |---|---|---|---|---|---|
85
+ | 1 | memory | 17 routes | `.loki/memory/` | MTIME | Keep |
86
+ | 2 | tasks | 8 routes | `.loki/queue/{pending,in-progress,completed,failed}.json` | PUSH + MTIME | Keep |
87
+ | 3 | council | 8 routes | `.loki/quality/reviews/<id>/` (`autonomy/run.sh:12736`) | MTIME | Keep |
88
+ | 4 | agents | 5 routes | `.loki/state/agents.json` | PUSH | Keep |
89
+ | 5 | cost | 3 routes | `.loki/metrics/budget.json` | PUSH on transition (`server.py:626-633`) | Keep |
90
+ | 6 | logs | 2 routes | `.loki/logs/` | MTIME | Keep |
91
+ | 7 | health | 3 routes | live process probe | ON-READ | Keep |
92
+ | 8 | events | 1 route | `.loki/events.jsonl` (`server.py:7080`) | MTIME + byte offset | Keep, extend |
93
+ | 9 | models | 3 routes | `.loki/state/model-override` | MTIME | Keep |
94
+ | 10 | providers | 1 route | `.loki/state/{provider,cli-provider}` | MTIME | Keep |
95
+ | 11 | **runs** | **5 routes exist** (`/api/v2/runs*`) | **SQL `Run` table, never written by the engine** | NONE | **Fix writer** |
96
+ | 12 | **receipts** | 0 routes | `.loki/proofs/<runId>/proof.json` (`loki-ts/src/runner/proof.ts:117`) | MTIME | **Add route** |
97
+ | 13 | **tests** | 0 routes | `.loki/quality/test-results.json` (`autonomy/run.sh:4781`), `.loki/verification/playwright-results.json` | MTIME | **Add route** |
98
+ | 14 | **artifacts** | 0 routes | `.loki/app-runner/state.json` (`run.sh:4196`), `first-preview.json` (`run.sh:4220`) | MTIME | **Add route** |
99
+ | 15 | **releases** | 0 routes | git tags (780 present), `VERSION`, `CHANGELOG.md` | ON-READ | **Add route** |
100
+ | - | **prompts** | 0 routes | **Review prompts only**: `.loki/quality/reviews/<id>/<reviewer>-prompt.txt` (`run.sh:14700`). Main-loop prompt is constructed in memory and never persisted. | MTIME (review only) | **Partial. Main loop NOT AVAILABLE** |
101
+
102
+ ### 1.3 The six domains, corrected
103
+
104
+ This is the table the brief asked for, with the brief's claim next to what the
105
+ code says.
106
+
107
+ | Domain | Brief says | Measured | Root cause | Work |
108
+ |---|---|---|---|---|
109
+ | runs | ZERO routes | **5 routes exist**, `api_v2.py` `/runs`, `/runs/{id}`, `/cancel`, `/replay`, `/timeline`; mounted `server.py:1010` | Only writer is `api_v2.py:461 create_run`. Nothing in `autonomy/` writes the `Run` table (grep over `autonomy/*.sh` finds only sqlalchemy dependency checks). Store is structurally empty for every real run. | **P0. Engine-side writer.** Not a new API. |
110
+ | receipts | ZERO routes | 0 routes. Confirmed. | Writer at `proof.ts:117` produces `.loki/proofs/<runId>/proof.json`. No instance in this worktree. | **P1. One read route.** |
111
+ | tests | ZERO routes | 0 routes. Confirmed. | Writer at `run.sh:4781` produces `.loki/quality/test-results.json`. No instance in this worktree. | **P1. One read route.** |
112
+ | artifacts | ZERO routes | 0 routes. Confirmed. | Writer at `run.sh:4196` produces `.loki/app-runner/state.json`. No instance in this worktree. | **P1. One read route.** |
113
+ | releases | ZERO routes | 0 routes. Confirmed. | Source is git itself (780 tags present, verified) plus `VERSION`. Only domain whose source is verified to exist. | **P2. One read route.** |
114
+ | prompts | ZERO routes | 0 routes. Confirmed. | Review prompts persisted at `run.sh:14700`. Main-loop prompt built in memory, never written. | **P2 for review prompts. Main loop = NOT AVAILABLE.** |
115
+
116
+ **The single most important line in this document:** runs is not a missing API.
117
+ It is a route reading an empty table. Building a second runs API would produce a
118
+ second empty surface.
119
+
120
+ ### 1.4 Contract rules for every new endpoint
121
+
122
+ Existing precedent to follow, not replace: `/api/events` reads
123
+ `.loki/events.jsonl` directly (`dashboard/server.py:7070-7092`), with an
124
+ existence check at `:7081` and a size cap at `:7089`. New file-backed domains
125
+ copy this shape.
126
+
127
+ Every endpoint returns an envelope:
128
+
129
+ ```
130
+ {
131
+ "data": <payload> | null,
132
+ "source": "<absolute-ish path or 'git' or 'computed'>",
133
+ "observed_at": "<RFC3339, server clock at read>",
134
+ "source_mtime": "<RFC3339>" | null,
135
+ "state": "ok" | "empty" | "stale" | "unavailable" | "error",
136
+ "detail": "<human-readable reason>" | null
137
+ }
138
+ ```
139
+
140
+ Rules, all binding:
141
+
142
+ 1. `source` is mandatory. A payload with no stated source is a defect.
143
+ 2. An unmeasured value is `null` with `state: "unavailable"`. **Never `0`**, and
144
+ never a plausible-looking sample. A zero cost and an unmeasured cost must not
145
+ render identically.
146
+ 3. `state` distinguishes four failure shapes that the current code collapses:
147
+ - `empty` - source exists, has no records (a run with no receipts yet)
148
+ - `unavailable` - source file absent (engine never ran)
149
+ - `stale` - source older than the domain's freshness budget
150
+ - `error` - source present but unparseable
151
+ 4. `source_mtime` is `null` only for ON-READ domains.
152
+ 5. HTTP status stays 200 for `empty` / `stale`, 503 for `unavailable`, 500 for
153
+ `error`. The UI must be able to distinguish these without parsing prose.
154
+
155
+ The existing 33 honesty markers in `server.py`
156
+ (`grep -oaE '"(unknown|unavailable|not_measured|UNKNOWN|not available)"' | wc -l`)
157
+ are the seed of this convention. The brief's "54" is not reproducible by that
158
+ command; I report 33 and the command that produced it. Either way, the markers
159
+ are ad hoc string literals today. The envelope makes the convention
160
+ machine-checkable instead of a naming habit.
161
+
162
+ ## 2. REALTIME MODEL
163
+
164
+ ### 2.1 What exists
165
+
166
+ - `/ws` at `dashboard/server.py:2499`. Auth by query parameter (`:2523-2540`)
167
+ because browsers cannot set headers on WS upgrade.
168
+ - `ConnectionManager` at `:478`. `MAX_CONNECTIONS` default 100 (`:482`).
169
+ - `broadcast` at `:512` fans out concurrently with a per-client
170
+ `SEND_TIMEOUT_SECONDS = 5.0` (`:510`) and **drops** any client that times out
171
+ (`:535-539`). Backpressure is already handled by disconnection, not buffering.
172
+ - Liveness: server pings after 30s of silence, closes after 2 missed pongs
173
+ (`:2554-2588`).
174
+ - `_push_loki_state_loop` at `:596`. Reads `dashboard-state.json`; broadcasts
175
+ only when mtime changed (`:670-672`). Budget (`:626`) and trust (`:646`)
176
+ transitions bypass the mtime gate.
177
+ - Other realtime surfaces: `/ws/collab`, `/api/managed/events`,
178
+ `/api/health/processes`.
179
+
180
+ Built artifact realtime constructs: 9
181
+ (`grep -oaE "new WebSocket|EventSource|setInterval" dashboard/static/index.html | wc -l`).
182
+ The brief's "34 realtime constructs" is not reproducible by any pattern I tried;
183
+ I mark that figure **not independently verified** rather than repeat it.
184
+
185
+ ### 2.2 The three gaps
186
+
187
+ **Gap 1 - silence is ambiguous.** `if mtime != last_mtime` (`:670`) means an
188
+ unchanged file produces no broadcast at all. A client cannot distinguish
189
+ "nothing changed" from "the writer died" from "my socket is half-open." This is
190
+ the exact defect the staleness requirement exists to prevent.
191
+
192
+ *Design answer:* every broadcast payload carries `source_mtime` and
193
+ `observed_at`. The client renders **age**, not just value. Past a per-domain
194
+ budget the panel flips to STALE. A panel showing a number with no age is
195
+ non-compliant. Silence is never evidence of freshness.
196
+
197
+ **Gap 2 - reconnect is blind.** The `connected` message (`:2547`) carries no
198
+ state snapshot. A reconnecting client waits for the next mtime change, which on
199
+ an idle run may never come.
200
+
201
+ *Design answer:* reconnect is REST-first. On open, the client fetches the
202
+ snapshot for each mounted panel over HTTP, then applies WS deltas on top. The
203
+ socket is an invalidation channel, not the source of truth. This also means a
204
+ total WS failure degrades to polling rather than to a blank dashboard.
205
+
206
+ **Gap 3 - no ordering or gap detection.** `broadcast` sends bare dicts with no
207
+ sequence number. A client cannot detect a dropped message, and since slow
208
+ clients are dropped mid-fan-out (`:535-539`), drops are an expected condition,
209
+ not an edge case.
210
+
211
+ *Design answer:* a monotonic `seq` per connection on every broadcast. Client
212
+ tracks last seen; on a gap it re-fetches the REST snapshot rather than trusting
213
+ its accumulated state. Ordering is per-connection and total; there is no
214
+ cross-domain ordering guarantee and the design does not pretend to one.
215
+
216
+ ### 2.3 Staleness budgets
217
+
218
+ | Domain class | Budget | Rationale |
219
+ |---|---|---|
220
+ | Active run state (tasks, agents) | 6s | Push loop runs at 2s; 3 missed cycles |
221
+ | Cost / budget | 60s | Transition-pushed, not periodic |
222
+ | Receipts, tests, artifacts | 120s | Written at phase boundaries, not continuously |
223
+ | Releases | 1h | Git tags change on release only |
224
+ | Health | 15s | ON-READ, staleness means the probe itself is failing |
225
+
226
+ On disconnect: panels do not clear and do not freeze silently. The header shows
227
+ a single global disconnected state, every panel dims and shows its last
228
+ `observed_at` age. Old numbers are never presented as current.
229
+
230
+ ## 3. INFORMATION ARCHITECTURE
231
+
232
+ The bundle is 779,725 bytes raw, 150,341 gzipped
233
+ (`dashboard/server.py:982`, `GZipMiddleware(minimum_size=1024)` at `:1001`).
234
+ gzip stays; nothing in this design touches it.
235
+
236
+ **Progressive disclosure here cannot mean code-splitting.**
237
+ `dashboard-ui/scripts/build-standalone.js` produces a single self-contained
238
+ file with zero runtime dependencies, written to both `dashboard/static/` and
239
+ `dist/`. That property is why the dashboard works offline and inside Docker
240
+ without a CDN. Splitting the bundle would trade a real capability for a
241
+ transfer-size win that gzip has already mostly taken (5.2x). So disclosure is
242
+ implemented as **deferred render and deferred fetch**, not deferred download.
243
+ The bytes arrive once; the work does not.
244
+
245
+ Concretely: a panel below the fold does not fetch, does not subscribe, and does
246
+ not render until it is revealed. The cost being managed is request fan-out and
247
+ DOM work against a 12k-line server, not kilobytes.
248
+
249
+ **First paint (new user, zero config).** One question answered: is anything
250
+ running, and is it healthy? Run status, current phase, budget, and a single
251
+ verdict tile. If no run has ever happened, this is an empty state that says so
252
+ and links to `loki start`. It must not render zeros.
253
+
254
+ **One click deep.** Task board, council reviews, cost breakdown, test results,
255
+ receipts, logs. Each fetches on reveal.
256
+
257
+ **Expert only.** Memory browser and graph, prompt optimizer, migration
258
+ dashboard, audit viewer, tenant switcher, raw event stream. These are the
259
+ heaviest panels and the least often needed; they stay behind explicit
260
+ navigation.
261
+
262
+ ## 4. MIGRATION
263
+
264
+ Measured on the shipped artifact with `grep -a`:
265
+
266
+ | Property | Count | Command |
267
+ |---|---|---|
268
+ | Custom elements registered | 39 | `grep -oa "customElements.define" dashboard/static/index.html \| wc -l` |
269
+ | `aria-*` attributes | 100 | `grep -oa "aria-[a-z]*=" ... \| wc -l` |
270
+ | `aria-` occurrences | 113 | `grep -oa "aria-" ... \| wc -l` |
271
+ | `role="` | 38 | `grep -oa 'role="' ... \| wc -l` |
272
+ | `@media` breakpoints | 13 | `grep -oa "@media" ... \| wc -l` |
273
+ | Component source files | 43 | `ls dashboard-ui/components \| grep -v vendor \| wc -l` |
274
+ | Registrations in source | 44 | `grep -roha "customElements.define" components core \| wc -l` |
275
+
276
+ The brief's 113 aria and 41 role are consistent with mine within pattern choice
277
+ (113 counts the bare `aria-` prefix, 100 counts full attributes). The component
278
+ figure is 43 files carrying 44 registrations, not 46.
279
+
280
+ **One gap worth noting:** source registers 44 custom elements, the shipped
281
+ artifact registers 39. Five components exist in `dashboard-ui/` and do not reach
282
+ the browser. This is out of scope for this document and I did not chase it, but
283
+ it is recorded here because it means the source tree is not a reliable proxy for
284
+ what ships. All acceptance criteria in section 5 therefore measure the **built
285
+ artifact**, not source.
286
+
287
+ ### Decision: keep
288
+
289
+ **Nothing is deleted. Nothing is rewritten.** The justification the brief asked
290
+ for cuts the other way: the accessibility work is real and shipped (100 aria
291
+ attributes, 38 roles across 39 elements), the zero-dependency build is a genuine
292
+ operational asset, and the responsive work exists. A rewrite would spend its
293
+ entire budget re-earning properties the codebase already has, while the actual
294
+ defect (no data) stayed untouched. A dashboard that is beautifully rebuilt and
295
+ still shows nothing real is a worse outcome than the current one.
296
+
297
+ | Component set | Action | Reason |
298
+ |---|---|---|
299
+ | All 39 registered elements | **Keep** | Working, accessible, zero-dependency |
300
+ | 138-156 existing routes | **Keep** | Working. No breaking changes |
301
+ | `/api/v2/runs*` (5 routes) | **Keep, fix writer** | Route is correct; store is empty |
302
+ | Envelope fields on responses | **Add, additive** | New optional keys; existing clients unaffected |
303
+ | `seq` on WS broadcasts | **Add, additive** | Unknown keys are ignored by current clients |
304
+ | Panels for the 5 new domains | **Add** | New elements alongside existing ones |
305
+ | gzip middleware | **Keep untouched** | 5.2x on the wire |
306
+
307
+ Every change is additive. The migration has no step at which the current
308
+ dashboard is worse than it is today.
309
+
310
+ ### Ranked work
311
+
312
+ | Rank | Work | Domain | Why first |
313
+ |---|---|---|---|
314
+ | P0 | Engine appends run lifecycle records at phase boundaries | runs | Unblocks 5 existing routes and the run-manager component. Highest value per unit of work in the entire plan |
315
+ | P1 | Read route over `.loki/proofs/<runId>/proof.json` | receipts | Trust core; writer already exists |
316
+ | P1 | Read route over `.loki/quality/test-results.json` | tests | Writer already exists |
317
+ | P1 | Read route over `.loki/app-runner/state.json` | artifacts | Writer already exists; drives preview |
318
+ | P2 | Read route over git tags + `VERSION` | releases | ON-READ, no writer needed, source verified present |
319
+ | P2 | Read route over `.loki/quality/reviews/<id>/*-prompt.txt` | prompts (review only) | Partial coverage, honestly labelled |
320
+ | P3 | Envelope + `seq` + staleness rendering | all | Correctness of everything above |
321
+
322
+ #### Why P0 needs a writer at all (alternative considered and rejected)
323
+
324
+ The cheaper option would be a read route over existing files, matching every P1
325
+ item and requiring no engine change. It was checked and does not work:
326
+
327
+ - `.loki/state/orchestrator.json` holds **current** phase and cumulative metrics
328
+ only (`autonomy/run.sh:6225-6226`, read at `:6438-6439`). It is overwritten in
329
+ place; there is no per-run history.
330
+ - `.loki/dashboard-state.json` is derived from that same file
331
+ (`autonomy/run.sh:6543-6556`), so it inherits the same limitation.
332
+ - `.loki/sessions/<id>/` contains only `loki.pgid` and `loki.pid`. Process
333
+ identity, not lifecycle.
334
+ - `.loki/checkout-runs/` contains a single numeric directory and is not a
335
+ general run store.
336
+
337
+ So current-run state is well served by files, and **run history has no file
338
+ source**. A history view cannot be built by reading what exists. That is why
339
+ this one item is a writer and everything else is a route.
340
+
341
+ Two constraints on that writer, both to bound R1. It appends at **phase
342
+ boundaries only**, never inside the iteration inner loop, so the hot path is
343
+ untouched. And it is **append-only**, so a crashed run leaves a truncated record
344
+ rather than a corrupt one. If phase-boundary granularity later proves too
345
+ coarse, the upgrade path is a finer trigger, not a different store.
346
+
347
+ Main-loop prompt capture is **not** on this list. It has no source of truth, and
348
+ inventing one is a change to the engine's hot path for a dashboard feature. It
349
+ renders NOT AVAILABLE.
350
+
351
+ ## 5. RISK REGISTER
352
+
353
+ | # | Risk | Evidence | Mitigation | Detection |
354
+ |---|---|---|---|---|
355
+ | R1 | Run writer double-writes or races the engine hot path | `run.sh` is 20k lines; state writes are scattered | Writer is append-only at phase boundaries; never in the iteration inner loop | Run duration before/after must not regress beyond noise |
356
+ | R2 | New file reads block the event loop on a large file | `/api/events` already caps reads (`server.py:7089`); learning aggregation reads up to 10MB (`:6904`) | Same size cap and existence check as the events precedent | Response time per new endpoint |
357
+ | R3 | `state: "empty"` renders as zero in a panel | Current code has 33 ad hoc markers, no shared convention | Envelope is mandatory; panels bind to `state` before `data` | A panel rendering a number while `state != "ok"` is a test failure |
358
+ | R4 | Slow-client drops read as data loss | `broadcast` drops on 5s timeout (`:535-539`) | `seq` gap triggers REST re-fetch | Client-side gap counter |
359
+ | R5 | 100-connection ceiling reached | `MAX_CONNECTIONS` (`:482`) | Already returns close code 1013; documented, not raised speculatively | Rejected-connection count |
360
+ | R6 | Bundle grows past the point gzip saves it | 779KB raw / 150KB gzipped today | Deferred render, not new dependencies. Zero-dependency property is a hard constraint | Raw and gzipped size per build |
361
+ | R7 | Additive envelope breaks a current consumer | Unknown | Fields are added, never renamed or removed | Existing route tests must pass unchanged |
362
+ | R8 | Measurement error is mistaken for a defect | Happened during this analysis: `grep` without `-a` reported 0 aria in a file with 113 | All artifact measurements use `grep -a`; empty results are treated as absent measurements | Any claim of "zero X" requires a positive control |
363
+
364
+ ## ACCEPTANCE MATRIX
365
+
366
+ Every row is measurable by a command. No aspirational entries.
367
+
368
+ | # | Criterion | Measurement | Pass condition |
369
+ |---|---|---|---|
370
+ | A1 | No route regression | `grep -cE '^@app\.(get\|post\|put\|delete\|patch\|websocket)' dashboard/server.py` | >= 165 |
371
+ | A2 | api_v2 routes intact | `grep -cE '@router\.(get\|post\|put\|delete)' dashboard/api_v2.py` | >= 24 |
372
+ | A3 | Accessibility not regressed | `grep -oa "aria-" dashboard/static/index.html \| wc -l` | >= 113 |
373
+ | A4 | Roles not regressed | `grep -oa 'role="' dashboard/static/index.html \| wc -l` | >= 38 |
374
+ | A5 | Custom elements not regressed | `grep -oa "customElements.define" dashboard/static/index.html \| wc -l` | >= 39 |
375
+ | A6 | Breakpoints not regressed | `grep -oa "@media" dashboard/static/index.html \| wc -l` | >= 13 |
376
+ | A7 | gzip still active | `grep -c "GZipMiddleware" dashboard/server.py` | >= 1 |
377
+ | A8 | Zero runtime dependencies | Built artifact contains no external script/CDN src | 0 external hosts |
378
+ | A9 | Runs store is written by the engine | Drive the lifecycle writer directly (source `run.sh`, call the phase-boundary function), then `GET /api/v2/runs`. No provider call, no spend | Returns >= 1 record with a real run id |
379
+ | A10 | Every new endpoint states source | Each new route's response | `source` key present and non-empty |
380
+ | A11 | Every new endpoint states freshness | Each new route's response | `observed_at` present; `source_mtime` present or explicitly null |
381
+ | A12 | Every new endpoint has an error state | Delete or corrupt the source file, call the route | Returns `state` in {`unavailable`,`error`}, never a 200 with zeros |
382
+ | A13 | No fabricated zeros | Call each new route with no `.loki/` present | No numeric field is `0`; all are null with `state: "unavailable"` |
383
+ | A14 | Staleness is visible | Freeze the source file past its budget | Panel shows STALE and an age |
384
+ | A15 | Disconnect is visible | Kill the WS server with the UI open | Every panel dims and shows last-updated age within 10s |
385
+ | A16 | Reconnect recovers without a push | Reconnect while source files are unchanged | Panels repopulate from REST |
386
+ | A17 | Gap detection works | Drop a broadcast | Client detects `seq` gap and re-fetches |
387
+ | A18 | Existing tests pass unchanged | Existing dashboard test suite | 0 failures, 0 modified assertions |
388
+
389
+ ### Not accepted as criteria
390
+
391
+ Deliberately excluded because they are not measurable as stated: "feels fast",
392
+ "enterprise-grade", "near-realtime" without a number, "clean UI", and any
393
+ coverage percentage (coverage is not measured in this release per
394
+ `skills/quality-gates.md`).
395
+
396
+ ## Appendix: contradictions to the brief
397
+
398
+ 1. **runs has 5 routes, not zero.** `dashboard/api_v2.py` `/runs`, `/runs/{id}`,
399
+ `/runs/{id}/cancel`, `/runs/{id}/replay`, `/runs/{id}/timeline`, mounted at
400
+ `dashboard/server.py:1010-1011`. The brief's route grep missed the `api_v2`
401
+ router because it uses `@router.` not `@app.`. The real defect is that the
402
+ only writer is `api_v2.py:461`; the engine never populates the table.
403
+ 2. **prompts is partially available.** Review prompts are persisted at
404
+ `autonomy/run.sh:14700` under `.loki/quality/reviews/<id>/<reviewer>-prompt.txt`.
405
+ Only the main-loop prompt is unavailable.
406
+ 3. **The other four domains have writers in code.** receipts
407
+ (`loki-ts/src/runner/proof.ts:117`), tests (`autonomy/run.sh:4781`),
408
+ artifacts (`autonomy/run.sh:4196`), releases (git, 780 tags verified
409
+ present). "No API" is true; "nothing produces this data" is false. This is
410
+ what changes the work from backend construction to route exposure. Note the
411
+ limit of the claim: for the first three, the writer is verified in source and
412
+ no output instance exists in this worktree.
413
+ 4. **Route count is 165 decorators / 156 unique method+path / 140 unique paths**
414
+ in `server.py`, plus 24 in `api_v2.py`. Not 138.
415
+ 5. **Honesty markers measure 33**, not 54, by
416
+ `grep -oaE '"(unknown|unavailable|not_measured|UNKNOWN|not available)"'`.
417
+ 6. **"34 realtime constructs" is not reproducible.** I measure 9 in the built
418
+ artifact. Marked not independently verified rather than repeated.
419
+ 7. **Component count is 43 files / 44 source registrations / 39 shipped
420
+ elements**, not 46/39. The source-to-shipped gap of 5 is unexplained and
421
+ out of scope here.
422
+ 8. The brief's aria (113) and breakpoint (13) figures **are confirmed**, but
423
+ only with `grep -a`. Without it the file greps as binary and reports zero.
package/docs/DEMOS.md CHANGED
@@ -1,26 +1,24 @@
1
1
  # Loki Mode -- Interactive Content Index
2
2
 
3
- Complete listing of all interactive HTML experiences in the Loki Mode repository. 18 files total across three categories.
3
+ Complete listing of all interactive HTML experiences in the Loki Mode repository. 13 files total across three categories.
4
4
 
5
- ## Walkthroughs (11 HTML files)
5
+ ## Walkthroughs (6 HTML files)
6
6
 
7
7
  Located in `docs/walkthrough/`.
8
8
 
9
- | File | Title | Description | Size |
10
- |------|-------|-------------|------|
11
- | `index.html` | Build Your First App | Step-by-step tutorial from PRD to deployment | 43 KB |
12
- | `architecture.html` | System Architecture | Interactive SVG diagram of the full system | 34 KB |
13
- | `comparison.html` | Feature Comparison | Loki Mode vs competitor feature matrix | 26 KB |
14
- | `gallery.html` | Project Gallery | 24 example projects across categories | 37 KB |
15
- | `video-placeholder.html` | Build Process Animation | CSS-animated build walkthrough | 25 KB |
16
- | `build-demo.html` | Full Build Video Demo | 12-iteration build replay with timeline | 18 KB |
17
- | `ide-demo.html` | Purple Lab IDE Demo | Browser-based IDE mockup with editor and build panel | 17 KB |
18
- | `enterprise-demo.html` | Enterprise Features Demo | Audit logging, RBAC, security, compliance | 13 KB |
19
- | `provider-race-demo.html` | Multi-Provider Race Demo | Claude vs Codex vs Gemini animated race | 17 KB |
20
- | `memory-demo.html` | Memory System Demo | Three-layer memory with interactive exploration | 19 KB |
21
- | `dashboard-demo.html` | Dashboard Monitoring Demo | Real-time dashboard with live build log | 17 KB |
22
-
23
- **Hub page**: `docs/walkthrough/hub.html` links to all walkthroughs with categories and estimated durations.
9
+ | File | Title | Description |
10
+ |------|-------|-------------|
11
+ | `index.html` | Build Your First App | Step-by-step tutorial from PRD to deployment |
12
+ | `architecture.html` | System Architecture | Interactive SVG diagram of the full system |
13
+ | `comparison.html` | Feature Comparison | Loki Mode vs competitor feature matrix |
14
+ | `gallery.html` | Project Gallery | Example projects across categories |
15
+ | `video-placeholder.html` | Build Process Animation | CSS-animated build walkthrough |
16
+ | `hub.html` | Walkthrough Hub | Links to all walkthroughs with categories and estimated durations |
17
+
18
+ An earlier version of this page also listed `build-demo.html`, `ide-demo.html`,
19
+ `enterprise-demo.html`, `provider-race-demo.html`, `memory-demo.html`, and
20
+ `dashboard-demo.html`. None of those files exist in the repository; they have
21
+ been removed from the list rather than left as broken references.
24
22
 
25
23
  ## Demo Apps (6 HTML files)
26
24
 
@@ -43,12 +41,12 @@ Located in `examples/`. Each was generated by Loki Mode from a single prompt.
43
41
 
44
42
  ## Summary
45
43
 
46
- | Category | Count | Total Size |
47
- |----------|-------|------------|
48
- | Walkthroughs | 11 | 266 KB |
49
- | Demo Apps | 6 | 65 KB |
50
- | Product Website | 1 | 38 KB |
51
- | **Total** | **18** | **369 KB** |
44
+ | Category | Count |
45
+ |----------|-------|
46
+ | Walkthroughs | 6 |
47
+ | Demo Apps | 6 |
48
+ | Product Website | 1 |
49
+ | **Total** | **13** |
52
50
 
53
51
  ## Viewing Locally
54
52