loki-mode 9.8.1 → 9.12.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (79) hide show
  1. package/README.md +19 -14
  2. package/SKILL.md +3 -2
  3. package/VERSION +1 -1
  4. package/autonomy/loki +122 -1
  5. package/autonomy/run.sh +49 -2
  6. package/dashboard/__init__.py +1 -1
  7. package/dashboard/api_evidence.py +411 -0
  8. package/dashboard/api_operator.py +283 -0
  9. package/dashboard/api_phases.py +262 -0
  10. package/dashboard/api_releases.py +242 -0
  11. package/dashboard/api_runs.py +477 -0
  12. package/dashboard/api_tests.py +444 -0
  13. package/dashboard/api_v2.py +47 -1
  14. package/dashboard/server.py +54 -0
  15. package/dashboard/static/index.html +246 -135
  16. package/docs/ARCHITECTURE-OVERVIEW.md +5 -3
  17. package/docs/CAPABILITY-BACKLOG.md +53 -0
  18. package/docs/COMPARISON.md +2 -2
  19. package/docs/COMPETITIVE-ANALYSIS.md +1 -1
  20. package/docs/COMPETITIVE-SCORECARD.md +422 -0
  21. package/docs/DASHBOARD-9.12-EVIDENCE.md +97 -0
  22. package/docs/DASHBOARD-ARCHITECTURE.md +423 -0
  23. package/docs/DEMOS.md +21 -23
  24. package/docs/HANDOFF-2026-08-03.md +439 -0
  25. package/docs/INSTALLATION.md +17 -10
  26. package/docs/OUTCOME-FRONTIER.md +536 -0
  27. package/docs/PROMPT-ABLATION-RESULT.md +97 -0
  28. package/docs/TOOLS.md +800 -0
  29. package/docs/alternative-installations.md +2 -3
  30. package/docs/audit-logging.md +44 -35
  31. package/docs/authentication.md +13 -2
  32. package/docs/authorization.md +87 -81
  33. package/docs/git-workflow.md +6 -3
  34. package/docs/metrics.md +15 -16
  35. package/docs/network-security.md +16 -13
  36. package/docs/openclaw-integration.md +36 -556
  37. package/docs/show-hn-post.md +2 -2
  38. package/docs/siem-integration.md +39 -36
  39. package/loki-ts/dist/loki.js +18 -18
  40. package/mcp/__init__.py +1 -1
  41. package/package.json +1 -1
  42. package/plugins/loki-mode/.claude-plugin/plugin.json +1 -1
  43. package/references/confidence-routing.md +18 -1
  44. package/references/invariant-checks.md +13 -8
  45. package/references/magic-rarv-integration.md +0 -1
  46. package/references/multi-provider.md +27 -5
  47. package/skills/healing.md +4 -2
  48. package/tools/audit-docs.py +488 -0
  49. package/tools/baseline-pin.py +19 -1
  50. package/tools/calibration-audit.py +523 -0
  51. package/tools/ci-gate.py +19 -1
  52. package/tools/cost-forecast.py +344 -0
  53. package/tools/cost-guard.py +19 -1
  54. package/tools/cost-history.py +19 -1
  55. package/tools/cost-per-outcome.py +394 -0
  56. package/tools/estimate-run.py +19 -1
  57. package/tools/evidence-freshness.py +307 -0
  58. package/tools/gate-init.py +19 -1
  59. package/tools/gate-report.py +19 -1
  60. package/tools/gate-simulate.py +570 -0
  61. package/tools/gate-trend.py +354 -0
  62. package/tools/model-advisor.py +52 -1
  63. package/tools/policy-load.py +19 -1
  64. package/tools/prompt-cost.py +363 -0
  65. package/tools/prompt-diff.py +448 -0
  66. package/tools/prompt-lint.py +448 -0
  67. package/tools/receipt-bundle.py +72 -2
  68. package/tools/receipt-diff.py +19 -1
  69. package/tools/receipt-find.py +19 -1
  70. package/tools/receipt-stats.py +380 -0
  71. package/tools/receipt-timeline.py +478 -0
  72. package/tools/receipt-verify-batch.py +291 -0
  73. package/tools/run-replay.py +19 -1
  74. package/tools/signing-status.py +19 -1
  75. package/tools/token-guard.py +19 -1
  76. package/tools/token-tax.py +375 -0
  77. package/tools/tool-index.py +19 -1
  78. package/tools/verification-tax.py +277 -0
  79. package/tools/verify-chain.py +361 -0
@@ -0,0 +1,536 @@
1
+ # Outcome Frontier: held-out tasks with deterministic oracles
2
+
3
+ > ## NOTHING HERE HAS BEEN RUN
4
+ >
5
+ > No task in this document has been executed. No tool has been scored. No cost
6
+ > has been spent producing it. Every number below is either a pre-registered
7
+ > threshold or an estimate labelled as one. There are no results in this file,
8
+ > and any cell that looks like a result is a worked example of the FORMAT, not
9
+ > an observation. If you are looking for evidence, this document contains none
10
+ > by design - it is the thing you run to GET evidence.
11
+
12
+ ## What this is, and what governs it
13
+
14
+ This is the operational layer under section 2b of
15
+ [`docs/COMPETITIVE-SCORECARD.md`](./COMPETITIVE-SCORECARD.md). That section
16
+ defines the tier ladder T0-T5 and the nine scored dimensions. This document does
17
+ not restate either; read 2b first.
18
+
19
+ One correction to the brief that commissioned this file, because a document
20
+ about honest measurement cannot open with an unverified citation. There is no
21
+ heading named `2b-FROZEN` anywhere in this repository (verified:
22
+ `grep -rn "2b-FROZEN" . --include='*.md'` returns nothing at the current
23
+ worktree commit). The frozen status block is the **"Status separation"** table
24
+ at the end of section 2b, which sorts every claim into Measured capability /
25
+ Planned experiments / Aspirational. That table governs this document, and this
26
+ document lands entirely in its **Planned experiments** row.
27
+
28
+ What 2b leaves open is the thing that decides whether the exercise is worth
29
+ running at all: *who decides whether a task succeeded.* 2b names dimensions. It
30
+ does not name oracles. This file supplies oracles for three of the six tiers and
31
+ is honest about which dimensions have no oracle and never will.
32
+
33
+ **Mapping to 2b tiers.** Three tasks, one per required shape:
34
+
35
+ | This doc | 2b tier | Shape |
36
+ |---|---|---|
37
+ | TASK-B | T1 | Brownfield change in an unfamiliar repo |
38
+ | TASK-M | T2 | Multi-file migration with real coupling |
39
+ | TASK-S | T5 | Scientific / research task - **PLAN ONLY** |
40
+
41
+ ## The property that decides whether any of this is worth running
42
+
43
+ An agent benchmark is worthless when the thing being measured also decides
44
+ whether it passed. This has three failure modes and all three are common:
45
+
46
+ 1. **Self-grading.** The agent reports its own success. The report is an output
47
+ of the system under test, so it measures the reporter, not the work.
48
+ 2. **LLM-judge-as-oracle.** A model scores free text against a rubric. This is
49
+ not deterministic, is not reproducible across model versions, and correlates
50
+ with verbosity and formatting. Dressing it as an oracle is the single most
51
+ common way a benchmark becomes a story.
52
+ 3. **In-context oracle.** The test the agent is scored against was visible in
53
+ the repo it was working in. The agent can pass by reading the answer, and the
54
+ benchmark cannot tell that apart from solving the problem.
55
+
56
+ Every oracle below is decidable by a process that is not the agent, not a
57
+ language model, and not visible to the agent during the run. Where a dimension
58
+ genuinely requires human judgement, it is listed in
59
+ [NOT DETERMINISTIC](#dimensions-that-are-not-deterministic-and-are-not-scored)
60
+ and is **not scored**. It is not converted into an LLM judge.
61
+
62
+ ---
63
+
64
+ ## TASK-B (T1): brownfield change in an unfamiliar repo
65
+
66
+ **Shape.** A bug fix or small feature inside a real, mature open-source
67
+ repository that the agent has never been given context for. Not a greenfield
68
+ build. The value being measured is comprehension of code someone else wrote.
69
+
70
+ **Instance selection (mechanical, not curated).** From a public repo pinned at a
71
+ specific commit SHA, select closed issues that (a) have a linked merge commit,
72
+ (b) whose merge commit modifies at least one non-test source file AND at least
73
+ one test file, (c) were closed AFTER the knowledge cutoff of every model under
74
+ test. Sort candidates by issue number and take the first N that satisfy the
75
+ filter. The person tuning the harness does not choose which issues are in the
76
+ set - the filter and the sort order do. See
77
+ [Held-out discipline](#held-out-discipline-mandatory).
78
+
79
+ **What the agent gets.** The repo checked out at the PARENT of the fix commit,
80
+ plus the issue title and body as written by the original reporter. Nothing else.
81
+ The issue text is not rewritten, hinted, or clarified - a rewritten issue is a
82
+ hint, and the whole point of brownfield is that real requests are underspecified.
83
+
84
+ **What is withheld.** The test files added or modified by the fix commit are
85
+ removed from the working tree the agent receives. This is load-bearing and must
86
+ be verified per instance before the run: if the oracle test is present in the
87
+ tree, the agent can read the expected behaviour directly and the oracle measures
88
+ nothing. An instance where the test cannot be cleanly withheld (for example, the
89
+ fix modifies an existing test the agent needs for unrelated context) is
90
+ **dropped from the set and the drop is reported**, not silently patched around.
91
+
92
+ ### Oracle
93
+
94
+ Post-hoc, the withheld test files are restored from the fix commit and the
95
+ repo's own test runner is executed against the agent's tree.
96
+
97
+ - **PASS**: every restored test passes AND the repo's pre-existing suite still
98
+ passes at its pinned baseline (no regressions).
99
+ - **PARTIAL**: restored tests pass but the pre-existing suite regresses, or a
100
+ strict subset of restored tests passes.
101
+ - **FAIL**: anything else, including a tree that does not build.
102
+
103
+ This is deterministic, runs offline, costs nothing beyond CPU, and is decided by
104
+ the upstream project's own test code written by the upstream maintainers before
105
+ this benchmark existed.
106
+
107
+ ### What this oracle CANNOT catch
108
+
109
+ - **A correct fix that the upstream test does not cover.** Upstream tests encode
110
+ the fix the maintainer actually wrote. A different, equally valid fix can fail
111
+ them. This biases toward the upstream solution and against novel ones.
112
+ - **A fix that passes by coincidence.** Tests can pass on a change that is right
113
+ for the wrong reason and will break on the next input.
114
+ - **Code quality of the diff.** A 400-line change and a 3-line change that both
115
+ pass score identically. Maintainability is not measured here (see NOT
116
+ DETERMINISTIC).
117
+ - **Overfitting to the test's literal assertions** if any part of the test
118
+ leaked into context through a stack trace, CI config, or changelog.
119
+ - **Whether the agent understood the issue** versus pattern-matched a similar
120
+ fix elsewhere in the repo.
121
+
122
+ ---
123
+
124
+ ## TASK-M (T2): multi-file migration with real coupling
125
+
126
+ **Shape.** Migrate every call site of an internal API to a replacement API
127
+ across a codebase where the call sites differ structurally - some in loops, some
128
+ behind conditionals, some with the old signature's argument order, some already
129
+ partly migrated. Coupling is the point: a task whose files are independent is N
130
+ single-file tasks wearing a trenchcoat.
131
+
132
+ **Instance selection.** Derived from a real migration commit in a pinned public
133
+ repo (same cutoff filter as TASK-B), or from a synthetic-but-fixed corpus
134
+ generated once, frozen, and stored outside this repo. If synthetic, the
135
+ generator's seed is recorded and the corpus is never regenerated between tools -
136
+ a regenerated corpus is a different experiment.
137
+
138
+ **What the agent gets.** The repo at the pre-migration commit, plus the
139
+ migration instruction as a specification of the NEW API's contract (signature,
140
+ semantics, deprecation reason). Not a list of the files to change - finding them
141
+ is the task.
142
+
143
+ ### Oracle
144
+
145
+ Three independent deterministic checks, all run post-hoc by the harness. All
146
+ three must hold for PASS.
147
+
148
+ 1. **Type checker / compiler.** The project's own `tsc --noEmit`, `mypy`, or
149
+ equivalent at its pinned config. Exit code only. Catches signature
150
+ mismatches and broken call sites.
151
+ 2. **Held-out integration tests.** Tests exercising the migrated behaviour
152
+ end-to-end, withheld from the agent's tree exactly as in TASK-B, restored and
153
+ run afterward. These assert BEHAVIOUR is preserved across the migration, not
154
+ just that it compiles.
155
+ 3. **Zero-residual-call-site invariant.** An AST query (not a grep - a grep
156
+ matches strings in comments and strings) asserting that zero call sites of
157
+ the old API remain reachable in source.
158
+
159
+ **The invariant must not be satisfiable by deletion.** Check 3 alone is trivially
160
+ passed by deleting every call site. It is only meaningful in conjunction with
161
+ check 2, which fails if the behaviour those call sites provided has disappeared.
162
+ Additionally the harness asserts a **call-site count floor**: the number of NEW
163
+ API call sites must be at least the number of OLD API call sites that existed
164
+ pre-migration. **The consolidation allowance is 0.** An agent may still
165
+ legitimately merge two call sites into one, but each such consolidation is
166
+ reported individually and justified against the behaviour check; a justified
167
+ consolidation does not count against the floor, and an unreported one fails it.
168
+ The allowance is fixed at 0 here rather than left as a threshold to be set
169
+ later, because an unspecified threshold in a pre-registration document is the
170
+ exact defect the frozen-rubric rule below names: a parameter set after seeing
171
+ results is a parameter chosen because of them. Stating this relationship is the
172
+ entire reason check 3 is not reported on its own.
173
+
174
+ - **PASS**: all three checks hold.
175
+ - **PARTIAL**: type check and invariant hold, integration tests partly fail.
176
+ - **FAIL**: type check fails, or call sites remain, or the count floor is
177
+ breached.
178
+
179
+ ### What this oracle CANNOT catch
180
+
181
+ - **Semantically wrong migrations that are behaviourally equivalent on the
182
+ tested paths.** Untested paths are unmeasured.
183
+ - **Consolidation that is technically legal but wrong** - the count floor's
184
+ allowance is a judgement call frozen in advance, and a clever wrong answer can
185
+ sit inside it.
186
+ - **Dynamic call sites.** Reflection, string-keyed dispatch, and
187
+ `getattr`-style access are invisible to an AST query. Instances that use them
188
+ should be excluded at selection time, and the exclusion reported.
189
+ - **Migration quality**: whether the resulting code is idiomatic, whether the
190
+ diff is reviewable, whether a human would merge it.
191
+ - **Partial credit for a correct-but-incomplete migration** that stopped at a
192
+ hard file. The oracle is binary per check; PARTIAL is coarse.
193
+
194
+ ---
195
+
196
+ ## TASK-S (T5): scientific / research task
197
+
198
+ > **THIS ONE IS A PLAN, NOT AN OPERATIONALIZED TASK.**
199
+ >
200
+ > Stated plainly and without hedging: TASK-B and TASK-M can be executed as
201
+ > written by someone who has not read anything else. TASK-S cannot. It requires
202
+ > a specific paper to be chosen, its artifact availability confirmed, and its
203
+ > headline number and tolerance transcribed before it is runnable. What follows
204
+ > is the design and the selection rule, not a runnable instance. Do not report
205
+ > TASK-S results alongside the other two as if they were equally rigorous.
206
+
207
+ **Shape.** Reproduce a published quantitative result. The agent is given a paper
208
+ (or its methods section) and the raw data or the data-acquisition instructions,
209
+ and must produce the reported figure through its own implementation.
210
+
211
+ **Why reproduction rather than open-ended research.** Open-ended research has no
212
+ deterministic oracle - novelty and usefulness require a domain expert, which is
213
+ why 2b lists them as needing a domain judge. Reproduction is the largest subset
214
+ of research work that IS decidable: the answer exists, it is a number, and the
215
+ agent has not been given it.
216
+
217
+ **Instance selection.** A paper with (a) a publicly available dataset, (b) a
218
+ headline numeric result stated with enough precision to compare against, (c)
219
+ publication date after every model's cutoff, or an obscure enough result that
220
+ memorization is implausible - and this second condition is weak, see RISKS. The
221
+ paper's own reference implementation, if one exists, is NOT given to the agent
222
+ and is used only as the differential comparison.
223
+
224
+ **What the agent gets.** The paper's methods section with the results section
225
+ removed, and the dataset. The number it is being scored against is not in its
226
+ context.
227
+
228
+ ### Oracle
229
+
230
+ Numeric reproduction within a pre-registered tolerance.
231
+
232
+ - The tolerance is **stated before the run**, derived from the paper's own
233
+ reported variance or confidence interval where one exists, or set at a fixed
234
+ relative tolerance recorded in the frozen rubric where one does not.
235
+ - **PASS**: the agent's produced figure falls within tolerance of the published
236
+ figure.
237
+ - **PARTIAL**: the agent produces a figure by a defensible method that falls
238
+ outside tolerance.
239
+ - **FAIL**: no figure produced, or the pipeline does not run.
240
+
241
+ Secondary deterministic check where a reference implementation exists:
242
+ differential comparison of outputs on identical input. This is a stronger signal
243
+ than the headline number alone, because a single scalar can be hit by accident.
244
+
245
+ ### What this oracle CANNOT catch
246
+
247
+ - **Memorization.** If the result is in the training data, reproduction measures
248
+ recall, not method. This is the dominant threat and it is not fully
249
+ mitigable - see RISKS.
250
+ - **Right number, wrong method.** A scalar within tolerance can be reached by a
251
+ pipeline that is wrong in ways that would diverge on any other dataset. The
252
+ differential check reduces this only where a reference implementation exists.
253
+ - **Whether the reproduction is a good scientific artifact** - reusable,
254
+ documented, correct in its statistics.
255
+ - **Novelty or usefulness.** Reproduction measures neither, by construction.
256
+ 2b's "Novelty/usefulness" dimension has no oracle here and is not scored.
257
+ - **The paper being wrong.** If the published number is itself an error, the
258
+ oracle rewards reproducing an error.
259
+
260
+ ---
261
+
262
+ ## Pre-registration
263
+
264
+ Everything in this section is frozen BEFORE the first run. This is the whole
265
+ difference between a benchmark and a story.
266
+
267
+ ### QUALITY
268
+
269
+ Scored solely by the per-task oracle above. Three values only: PASS / PARTIAL /
270
+ FAIL, with the per-task definitions as written. No aggregate quality score is
271
+ computed across tasks - three tasks cannot support an average, and a mean over
272
+ three ordinal values is not a measurement.
273
+
274
+ There is **no LLM-judged quality dimension**. If a quality question cannot be
275
+ answered by the oracle, it appears in the NOT DETERMINISTIC list and goes
276
+ unscored.
277
+
278
+ ### COST
279
+
280
+ - **Measured**: total USD spend attributable to the run.
281
+ - **Read from**: the per-run receipt's `cost.usd`, surfaced as `cost_usd`, which
282
+ is the convention `tools/cost-per-outcome.py` and `tools/receipt-stats.py`
283
+ already read (`tools/cost-per-outcome.py:151`,
284
+ `tools/receipt-stats.py:138`). This document adopts that convention rather
285
+ than inventing a second one.
286
+ - **UNKNOWN, never 0.** A run that recorded no cost did not cost zero dollars.
287
+ It reads UNKNOWN, it is excluded from every total, and the exclusion is
288
+ reported on the same row as the figure it was excluded from. This rule is
289
+ inherited verbatim from `tools/cost-per-outcome.py`, whose module docstring
290
+ states it as honesty rule 1 and enforces it by letting the MEASURED COUNT
291
+ decide UNKNOWN rather than the total. An unmeasured run is not evidence of
292
+ cheapness; it is an absent measurement.
293
+
294
+ **Cross-tool comparability - the load-bearing caveat.** The above reads Loki's
295
+ artifacts. Competing tools emit no receipt. An oracle that can only read our own
296
+ cost artifact silently advantages us, and a founder reading `$X vs UNKNOWN` will
297
+ read UNKNOWN as "free" or as "broken" when it means neither.
298
+
299
+ The pre-registered rule:
300
+
301
+ 1. Cost is captured uniformly where the provider exposes a per-run usage record
302
+ independent of the tool (provider-side billing or usage API, bracketed by the
303
+ harness around the run). That figure, not the tool's self-report, is the
304
+ comparable one.
305
+ 2. Where a tool provides no such record, its cost reads **UNKNOWN** for that run.
306
+ 3. **A measured figure and an UNKNOWN cannot be differenced, ratioed, or ranked
307
+ against each other.** A cost comparison is reported only across tools that
308
+ were all measured by the same route. If only one tool has a measurable cost,
309
+ there is no cost comparison in that experiment, and the section says so
310
+ instead of showing one column with a number and one column empty.
311
+
312
+ ### TIME
313
+
314
+ - **Measured**: wall clock from run start to run end.
315
+ - **Read from**: the tool's own emitted artifacts. For Loki, timestamps in
316
+ `.loki/events.jsonl` (flat `{"timestamp","type","data"}` schema, per
317
+ `autonomy/completion-council.sh:3893`), and where the run produces a preview,
318
+ `.loki/app-runner/first-preview.json`. For tools that emit no timestamped
319
+ artifact, an external timestamp bracket written by the harness immediately
320
+ before spawn and immediately after exit - written by the harness, not by the
321
+ tool.
322
+ - **NEVER `ps etime` on a supervising process.** This is not a stylistic
323
+ preference. This repo hit exactly this bug: a hung benchmark harness reported
324
+ 39 minutes for a 20-minute run because `etime` measured the WAITER's lifetime,
325
+ not the work. `pgrep -f` additionally matches the waiter itself. Any time
326
+ figure sourced from a process table is void and is re-measured, not adjusted.
327
+
328
+ **Cross-tool comparability applies to TIME exactly as it does to COST**, and for
329
+ a sharper reason. Our `events.jsonl` clock starts after the engine is already
330
+ up; a harness bracket around a competitor includes process spawn, client init,
331
+ and teardown. Differencing the two makes our number smaller for reasons that are
332
+ not speed - RISK 7 (harness asymmetry) expressing itself inside a scored
333
+ dimension. Therefore: **the comparable wall clock is the harness bracket,
334
+ applied uniformly to every tool INCLUDING ours.** The tool's own artifacts
335
+ (`.loki/events.jsonl`, `.loki/app-runner/first-preview.json`) are the
336
+ finer-grained internal breakdown - reported, never used for a cross-tool
337
+ difference. A harness-written timestamp around spawn and exit is not a
338
+ process-table read, so this satisfies the prohibition above rather than evading
339
+ it.
340
+
341
+ ### The frozen rubric
342
+
343
+ The scored dimensions are exactly: **QUALITY** (oracle verdict), **COST**,
344
+ **TIME**. Three. That is the entire rubric.
345
+
346
+ > **A dimension added after seeing results is a dimension chosen because of
347
+ > them.** If, after the first run, someone proposes a fourth dimension, that
348
+ > proposal is evidence about the results and not about the rubric. The honest
349
+ > handling is to run it as a NEW, separately-reported experiment against a
350
+ > re-frozen rubric, and to state in the report that the dimension was added
351
+ > post-hoc and on which observation. Silently widening the rubric converts a
352
+ > benchmark into a search for a framing in which we win.
353
+
354
+ ---
355
+
356
+ ## Held-out discipline (mandatory)
357
+
358
+ 1. **Task instances are not selected by whoever tunes the harness.** Selection
359
+ is by the mechanical filter stated per task - repo pinned at a SHA, filter
360
+ predicate, deterministic sort, first N. The filter is written down before the
361
+ candidate list is inspected. If an instance is dropped (unwithhold-able test,
362
+ dynamic dispatch, missing dataset), the drop and its reason are reported
363
+ alongside the results. An unreported drop is a selection effect.
364
+
365
+ 2. **One scored attempt per tool per task.** The first run is the run. A retry
366
+ after seeing the result is a DIFFERENT EXPERIMENT: it is reported separately,
367
+ labelled as a retry, with the observation that prompted it stated. It is
368
+ never averaged into the first attempt, and the first attempt is never
369
+ replaced. "Best of N" where N was chosen after seeing the failures is not a
370
+ measurement of capability; it is a measurement of how many tries the
371
+ experimenter was willing to fund.
372
+
373
+ 3. **Task text is NEVER committed to this repo.** A task stored in the
374
+ repository is a task inside the context of the thing being measured - any
375
+ agent running in this tree can read it, and every future model trained on
376
+ this repo has seen it. This document therefore contains task SHAPES,
377
+ SELECTION RULES, and ORACLE DEFINITIONS, and contains no task text, no repo
378
+ URL, no issue number, no paper title, and no expected value.
379
+
380
+ **Where they live instead**: task instances are stored outside this
381
+ repository and outside any directory an agent under test is given - a private
382
+ store, referenced from the run harness by identifier only. The run report
383
+ cites instances by opaque ID. Resolving an ID to its content is a manual step
384
+ performed by the person running the experiment. If an instance's text ever
385
+ appears in this repo, in a commit message, or in a run log committed here,
386
+ that instance is **burned** and is replaced, not reused.
387
+
388
+ ---
389
+
390
+ ## Calibration caveat (must not be softened)
391
+
392
+ `tools/calibration-audit.py` scores **AGREEMENT WITH THE COUNCIL MAJORITY, NOT
393
+ ACCURACY.** The council's outcome is mechanically derived from the votes
394
+ (`approve_count >= threshold`), so a voter's own prediction partially CAUSES the
395
+ label it is subsequently scored against. A voter scoring perfectly there may
396
+ simply be voting with the crowd. No artifact on disk records whether the council
397
+ was actually RIGHT, so ground-truth calibration is not computable from that
398
+ substrate at all.
399
+
400
+ Therefore, in this document and any report generated from it:
401
+
402
+ - Calibration is **a caveat, never an accuracy term and never a quality term.**
403
+ - Calibration is **never a scored dimension.** It is not in the frozen rubric
404
+ above and may not be added to it. A high calibration figure is not evidence
405
+ that any task was completed well, and must never be cited as though it were.
406
+
407
+ Factual note on availability: `tools/calibration-audit.py` is not present in
408
+ this worktree at the current commit. It exists on another branch (introduced in
409
+ commit `8c33d123`), whose own module docstring states this circularity in the
410
+ same terms. The caveat above stands at full strength regardless of which branch
411
+ the tool is on - it is a prohibition on how a class of number may be used, not a
412
+ description of a file.
413
+
414
+ ---
415
+
416
+ ## Dimensions that are NOT deterministic and are NOT scored
417
+
418
+ These are inherited from 2b's scored-dimension table, where they are defined.
419
+ They are listed here because this document's job is to say which of them have
420
+ oracles. **These do not.** They are marked NOT DETERMINISTIC and left unscored
421
+ rather than approximated with an LLM judge.
422
+
423
+ | Dimension (defined in 2b) | Why it has no oracle |
424
+ |---|---|
425
+ | Completion quality *beyond the oracle verdict* | 2b defines it as judged against intent. Intent is not a file. The oracle covers "does it pass"; "is it what was wanted" is human. |
426
+ | Maintainability | Requires reading the diff as a future maintainer. No checker distinguishes a clean fix from a working mess. |
427
+ | Intervention rate | Countable in principle, but what counts as an intervention is a judgement about whether a nudge was necessary or merely habitual. |
428
+ | Novelty / usefulness (T5) | 2b already states this needs a domain judge. Reproduction deliberately sidesteps it and therefore does not measure it. |
429
+ | User value ("would the requester ship it") | 2b's own framing: the only dimension that ends in a human. It ends in a human here too. |
430
+ | Recovery after induced failure | Requires deliberate fault injection not designed here; the fault set would be chosen by the harness tuner, which is the selection effect this document exists to avoid. |
431
+ | Verification tax | The time/token half is measurable; "how often it changed an outcome" requires knowing the counterfactual outcome, which is not observable from a single run per task. |
432
+
433
+ Being unscored is not the same as being unimportant. Several of these matter
434
+ more than the ones that are scored. They are excluded because a fabricated
435
+ measurement of an important thing is worse than an acknowledged gap.
436
+
437
+ ---
438
+
439
+ ## RISKS: what could make these results misleading even if executed perfectly
440
+
441
+ 1. **Training-data contamination.** The dominant threat. Public repos, their
442
+ issues, their fixes, and published papers with their numbers are all
443
+ plausibly in training data. Post-cutoff filtering is a mitigation, not a
444
+ solution: cutoffs are approximate, self-reported, and differ per tool, so a
445
+ date filter that holds one model out may not hold another. A tool that
446
+ memorized the fix and a tool that derived it produce identical oracle
447
+ verdicts. **This affects TASK-S most severely** and is why TASK-S results, if
448
+ ever produced, carry lower evidential weight than TASK-B or TASK-M.
449
+
450
+ 2. **n is far too small for a ranking.** Three task shapes, one instance each,
451
+ one attempt each. This can support existence claims ("tool X solved this
452
+ instance") and cannot support comparative claims ("tool X is better than tool
453
+ Y"). Any percentage computed over three tasks is theatre. Reporting a winner
454
+ from this design would be the central misuse.
455
+
456
+ 3. **Selection effects that survive the mechanical filter.** The filter is
457
+ mechanical but the CHOICE OF FILTER is not - which repo, which language,
458
+ which issue-shape predicate. A filter that happens to select tasks resembling
459
+ our own test corpus advantages us. Mitigation: fix the repo and filter before
460
+ inspecting candidates, and report the drop list.
461
+
462
+ 4. **Oracle gaming.** A held-out test suite is a target once the agent infers it
463
+ exists. Agents can write code that special-cases plausible test inputs, or
464
+ over-fit to behaviour implied by the issue text. TASK-M's residual-call-site
465
+ invariant is specifically gameable by deletion, which is why it is never
466
+ reported without the behaviour check and the count floor. Assume any oracle
467
+ stated publicly will eventually be optimized against.
468
+
469
+ 5. **Tool-version drift mid-experiment.** These CLIs ship frequently - this repo
470
+ itself releases hourly. A run of tool A on Monday and tool B on Wednesday
471
+ compares two different weeks, not two tools. Mitigation: pin and record the
472
+ exact version of every tool, and re-run the whole matrix if any version
473
+ changes mid-experiment. A partially-refreshed matrix is void.
474
+
475
+ 6. **Environment non-determinism.** Network flakiness, rate limits, provider-side
476
+ model routing, and nondeterministic sampling mean the same tool on the same
477
+ task can produce different outcomes. With one attempt per task, a single rate
478
+ limit is indistinguishable from a capability failure. Rate-limit and
479
+ infrastructure failures must be recorded as INFRASTRUCTURE, not FAIL, and
480
+ such a run is re-run - which is itself a deviation from one-attempt
481
+ discipline and must be reported as one.
482
+
483
+ 7. **Harness asymmetry.** We wrote the harness. Our tool's artifacts are
484
+ first-class to it; competitors are driven through whatever CLI surface they
485
+ expose. Effort spent making our own path work smoothly and not theirs is an
486
+ advantage that looks like capability. Mitigation: drive every tool through
487
+ its documented non-interactive entry point only, and report any tool for
488
+ which the harness needed tool-specific accommodation.
489
+
490
+ 8. **The oracle can be wrong.** Upstream tests can be flaky, papers can report
491
+ errors, and type checkers have config-dependent behaviour. A deterministic
492
+ oracle is reproducible, which is not the same as correct.
493
+
494
+ ---
495
+
496
+ ## COST-TO-RUN estimate
497
+
498
+ **This is an estimate, not a measurement.** No run has occurred; nothing here is
499
+ read from an artifact. It exists so the expenditure can be approved or declined
500
+ before rather than after.
501
+
502
+ **Assumptions** (each is a place the estimate can be wrong):
503
+
504
+ - 3 tasks x 1 attempt x the number of tools compared. At 4 tools that is 12
505
+ scored runs.
506
+ - Per-run agentic spend of roughly $2-$15 depending on tool, model tier, and
507
+ iteration count. The spread is wide because iteration count is the dominant
508
+ cost term and is exactly what the experiment is measuring - it cannot be
509
+ known in advance without circularity.
510
+ - Oracle execution (test runs, type checks, AST queries) is local CPU only:
511
+ effectively $0 in provider spend.
512
+ - TASK-S is excluded from this estimate because it is a plan, not a runnable
513
+ task. Adding it means adding dataset acquisition and compute of unknown size.
514
+ - Setup labour (instance selection, test-withholding verification, harness
515
+ wiring per tool) is human time and is NOT included in the dollar figure.
516
+
517
+ **Estimated provider spend**: roughly **$25 to $180** for a 4-tool, 2-task
518
+ (TASK-B and TASK-M) matrix. The upper bound is what to budget; the lower bound
519
+ is what to hope for.
520
+
521
+ **Estimated wall clock**: roughly **4 to 12 hours** of run time for the matrix
522
+ if run serially, plus an estimated **1 to 3 days of human setup** dominated by
523
+ instance selection and by verifying per instance that the oracle tests are
524
+ genuinely withheld. The setup is the expensive half and is not parallelizable by
525
+ adding tools.
526
+
527
+ **What the money buys.** Existence claims on two task shapes with
528
+ non-self-graded oracles, and a reusable harness. It does NOT buy a ranking - see
529
+ RISKS 2. A founder approving this should approve it as instrumentation, not as a
530
+ competitive result, and should decline it if a ranking is what is wanted, because
531
+ this design cannot produce one honestly.
532
+
533
+ ---
534
+
535
+ *Status per section 2b's Status separation table: **Planned experiments**. Zero
536
+ runs to date.*
@@ -0,0 +1,97 @@
1
+ # Prompt ablation: measured result
2
+
3
+ Date: 2026-08-03. Instrument: `benchmarks/run-prompt-ablation.sh`,
4
+ reporter `benchmarks/report-prompt-ablation.py`. Raw rows:
5
+ `benchmarks/results/prompt-ablation.jsonl`.
6
+
7
+ ## The question
8
+
9
+ Anthropic deleted roughly 80% of Claude Code's system prompt for Opus 5, on
10
+ the finding that instructions written to correct older models had become dead
11
+ weight, and that the model measured slightly MORE capable without them.
12
+
13
+ `LOKI_SIMPLE=1` strips the coaching half of our prompt: the RARV cycle, SDLC
14
+ phases, memory habits. Per-iteration state (which gate failed, self-heal
15
+ output, checklist status) is never touched, because that is information the
16
+ model cannot derive from anywhere else.
17
+
18
+ Prompt size fell 78% (about 1562 tokens per iteration), reproduced
19
+ independently on both the bash and Bun routes. That is a TOKEN measurement.
20
+ It says nothing about whether the stripped arm builds the same thing, faster
21
+ or slower, more or less reliably. This is that measurement.
22
+
23
+ ## Result
24
+
25
+ ```
26
+ trials recorded: full=3 simple=3
27
+
28
+ metric full (default) simple (LOKI_SIMPLE=1)
29
+ -----------------------------------------------------------
30
+ wall clock 8.2 min [7.3-8.7] n=3 7.6 min [6.2-8.0] n=3
31
+ iterations 1.0 [1.0-1.0] n=3 1.0 [1.0-1.0] n=3
32
+
33
+ full reliability: 3/3 completed
34
+ simple reliability: 3/3 completed
35
+
36
+ VERDICT: NO DIFFERENCE DEMONSTRATED -- the arms' ranges overlap
37
+ (full 7.3-8.7, simple 6.2-8.0), so the median gap is within
38
+ the run-to-run noise.
39
+ ```
40
+
41
+ ## What this means, stated carefully
42
+
43
+ **Speed: no difference demonstrated.** The simple arm's median is 0.6 min
44
+ lower, but the ranges overlap heavily. A gap smaller than the variance each
45
+ arm shows against itself is not a finding.
46
+
47
+ **Reliability: no difference demonstrated.** 3/3 completed in both arms, and
48
+ every run passed its acceptance check. Both arms finished in a single
49
+ iteration.
50
+
51
+ **Cost: the token saving is real and stands on its own.** 78% less prompt per
52
+ iteration is a deterministic, reproducible measurement that does not depend on
53
+ these trials at all.
54
+
55
+ **The honest summary: stripping the coaching prompt cost us nothing
56
+ measurable, and saved 78% of the prompt.** That is a weaker claim than "the
57
+ model is better without it" and a stronger one than "no effect" -- removing
58
+ 1562 tokens of instruction per iteration did not degrade speed, completion, or
59
+ acceptance across six real builds.
60
+
61
+ ## Why this was nearly reported wrong
62
+
63
+ After trial 1, the numbers were full 8.2 min against simple 6.2 min. That is a
64
+ 24% improvement, and it is exactly the shape of result that gets written into
65
+ a release note.
66
+
67
+ Trial 2 reversed it: full 7.3, simple 8.0.
68
+
69
+ One trial of a stochastic agent is indistinguishable from noise. The reporter
70
+ refuses to declare a winner when the ranges overlap, and refuses entirely
71
+ below n=3 per arm, which is why the instrument was built before the claim
72
+ rather than after it.
73
+
74
+ ## Limits, stated rather than left for a reader to discover
75
+
76
+ - **n=3 per arm.** Enough to see that the arms overlap; not enough to detect a
77
+ small real effect. A difference under roughly 20% would not be visible here.
78
+ - **One spec, one model.** A single small build on sonnet. Coaching may matter
79
+ more on a longer task, a weaker model, or a spec where the RARV structure is
80
+ load-bearing. This says nothing about those.
81
+ - **Both arms finished in one iteration**, so this never exercised the
82
+ iterate-and-recover path where per-iteration state matters most. That is the
83
+ case where stripping coaching is most likely to be safe and where the
84
+ remaining prompt does the most work, and it is untested.
85
+ - **Degraded providers are unaffected.** Codex and Aider take an earlier
86
+ return path, so the flag never reaches its gate there. Measured as a 0-byte
87
+ delta, not assumed.
88
+
89
+ ## Reproduce
90
+
91
+ ```bash
92
+ benchmarks/run-prompt-ablation.sh --trials 3 --model sonnet --max-iters 6
93
+ python3 benchmarks/report-prompt-ablation.py benchmarks/results/prompt-ablation.jsonl
94
+ ```
95
+
96
+ Both arms run in ONE engine copy; the only difference is the environment
97
+ variable read at prompt-assembly time.