loki-mode 9.8.1 → 9.12.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +19 -14
- package/SKILL.md +3 -2
- package/VERSION +1 -1
- package/autonomy/loki +122 -1
- package/autonomy/run.sh +49 -2
- package/dashboard/__init__.py +1 -1
- package/dashboard/api_evidence.py +411 -0
- package/dashboard/api_operator.py +283 -0
- package/dashboard/api_phases.py +262 -0
- package/dashboard/api_releases.py +242 -0
- package/dashboard/api_runs.py +477 -0
- package/dashboard/api_tests.py +444 -0
- package/dashboard/api_v2.py +47 -1
- package/dashboard/server.py +54 -0
- package/dashboard/static/index.html +246 -135
- package/docs/ARCHITECTURE-OVERVIEW.md +5 -3
- package/docs/CAPABILITY-BACKLOG.md +53 -0
- package/docs/COMPARISON.md +2 -2
- package/docs/COMPETITIVE-ANALYSIS.md +1 -1
- package/docs/COMPETITIVE-SCORECARD.md +422 -0
- package/docs/DASHBOARD-9.12-EVIDENCE.md +97 -0
- package/docs/DASHBOARD-ARCHITECTURE.md +423 -0
- package/docs/DEMOS.md +21 -23
- package/docs/HANDOFF-2026-08-03.md +439 -0
- package/docs/INSTALLATION.md +17 -10
- package/docs/OUTCOME-FRONTIER.md +536 -0
- package/docs/PROMPT-ABLATION-RESULT.md +97 -0
- package/docs/TOOLS.md +800 -0
- package/docs/alternative-installations.md +2 -3
- package/docs/audit-logging.md +44 -35
- package/docs/authentication.md +13 -2
- package/docs/authorization.md +87 -81
- package/docs/git-workflow.md +6 -3
- package/docs/metrics.md +15 -16
- package/docs/network-security.md +16 -13
- package/docs/openclaw-integration.md +36 -556
- package/docs/show-hn-post.md +2 -2
- package/docs/siem-integration.md +39 -36
- package/loki-ts/dist/loki.js +18 -18
- package/mcp/__init__.py +1 -1
- package/package.json +1 -1
- package/plugins/loki-mode/.claude-plugin/plugin.json +1 -1
- package/references/confidence-routing.md +18 -1
- package/references/invariant-checks.md +13 -8
- package/references/magic-rarv-integration.md +0 -1
- package/references/multi-provider.md +27 -5
- package/skills/healing.md +4 -2
- package/tools/audit-docs.py +488 -0
- package/tools/baseline-pin.py +19 -1
- package/tools/calibration-audit.py +523 -0
- package/tools/ci-gate.py +19 -1
- package/tools/cost-forecast.py +344 -0
- package/tools/cost-guard.py +19 -1
- package/tools/cost-history.py +19 -1
- package/tools/cost-per-outcome.py +394 -0
- package/tools/estimate-run.py +19 -1
- package/tools/evidence-freshness.py +307 -0
- package/tools/gate-init.py +19 -1
- package/tools/gate-report.py +19 -1
- package/tools/gate-simulate.py +570 -0
- package/tools/gate-trend.py +354 -0
- package/tools/model-advisor.py +52 -1
- package/tools/policy-load.py +19 -1
- package/tools/prompt-cost.py +363 -0
- package/tools/prompt-diff.py +448 -0
- package/tools/prompt-lint.py +448 -0
- package/tools/receipt-bundle.py +72 -2
- package/tools/receipt-diff.py +19 -1
- package/tools/receipt-find.py +19 -1
- package/tools/receipt-stats.py +380 -0
- package/tools/receipt-timeline.py +478 -0
- package/tools/receipt-verify-batch.py +291 -0
- package/tools/run-replay.py +19 -1
- package/tools/signing-status.py +19 -1
- package/tools/token-guard.py +19 -1
- package/tools/token-tax.py +375 -0
- package/tools/tool-index.py +19 -1
- package/tools/verification-tax.py +277 -0
- package/tools/verify-chain.py +361 -0
|
@@ -0,0 +1,536 @@
|
|
|
1
|
+
# Outcome Frontier: held-out tasks with deterministic oracles
|
|
2
|
+
|
|
3
|
+
> ## NOTHING HERE HAS BEEN RUN
|
|
4
|
+
>
|
|
5
|
+
> No task in this document has been executed. No tool has been scored. No cost
|
|
6
|
+
> has been spent producing it. Every number below is either a pre-registered
|
|
7
|
+
> threshold or an estimate labelled as one. There are no results in this file,
|
|
8
|
+
> and any cell that looks like a result is a worked example of the FORMAT, not
|
|
9
|
+
> an observation. If you are looking for evidence, this document contains none
|
|
10
|
+
> by design - it is the thing you run to GET evidence.
|
|
11
|
+
|
|
12
|
+
## What this is, and what governs it
|
|
13
|
+
|
|
14
|
+
This is the operational layer under section 2b of
|
|
15
|
+
[`docs/COMPETITIVE-SCORECARD.md`](./COMPETITIVE-SCORECARD.md). That section
|
|
16
|
+
defines the tier ladder T0-T5 and the nine scored dimensions. This document does
|
|
17
|
+
not restate either; read 2b first.
|
|
18
|
+
|
|
19
|
+
One correction to the brief that commissioned this file, because a document
|
|
20
|
+
about honest measurement cannot open with an unverified citation. There is no
|
|
21
|
+
heading named `2b-FROZEN` anywhere in this repository (verified:
|
|
22
|
+
`grep -rn "2b-FROZEN" . --include='*.md'` returns nothing at the current
|
|
23
|
+
worktree commit). The frozen status block is the **"Status separation"** table
|
|
24
|
+
at the end of section 2b, which sorts every claim into Measured capability /
|
|
25
|
+
Planned experiments / Aspirational. That table governs this document, and this
|
|
26
|
+
document lands entirely in its **Planned experiments** row.
|
|
27
|
+
|
|
28
|
+
What 2b leaves open is the thing that decides whether the exercise is worth
|
|
29
|
+
running at all: *who decides whether a task succeeded.* 2b names dimensions. It
|
|
30
|
+
does not name oracles. This file supplies oracles for three of the six tiers and
|
|
31
|
+
is honest about which dimensions have no oracle and never will.
|
|
32
|
+
|
|
33
|
+
**Mapping to 2b tiers.** Three tasks, one per required shape:
|
|
34
|
+
|
|
35
|
+
| This doc | 2b tier | Shape |
|
|
36
|
+
|---|---|---|
|
|
37
|
+
| TASK-B | T1 | Brownfield change in an unfamiliar repo |
|
|
38
|
+
| TASK-M | T2 | Multi-file migration with real coupling |
|
|
39
|
+
| TASK-S | T5 | Scientific / research task - **PLAN ONLY** |
|
|
40
|
+
|
|
41
|
+
## The property that decides whether any of this is worth running
|
|
42
|
+
|
|
43
|
+
An agent benchmark is worthless when the thing being measured also decides
|
|
44
|
+
whether it passed. This has three failure modes and all three are common:
|
|
45
|
+
|
|
46
|
+
1. **Self-grading.** The agent reports its own success. The report is an output
|
|
47
|
+
of the system under test, so it measures the reporter, not the work.
|
|
48
|
+
2. **LLM-judge-as-oracle.** A model scores free text against a rubric. This is
|
|
49
|
+
not deterministic, is not reproducible across model versions, and correlates
|
|
50
|
+
with verbosity and formatting. Dressing it as an oracle is the single most
|
|
51
|
+
common way a benchmark becomes a story.
|
|
52
|
+
3. **In-context oracle.** The test the agent is scored against was visible in
|
|
53
|
+
the repo it was working in. The agent can pass by reading the answer, and the
|
|
54
|
+
benchmark cannot tell that apart from solving the problem.
|
|
55
|
+
|
|
56
|
+
Every oracle below is decidable by a process that is not the agent, not a
|
|
57
|
+
language model, and not visible to the agent during the run. Where a dimension
|
|
58
|
+
genuinely requires human judgement, it is listed in
|
|
59
|
+
[NOT DETERMINISTIC](#dimensions-that-are-not-deterministic-and-are-not-scored)
|
|
60
|
+
and is **not scored**. It is not converted into an LLM judge.
|
|
61
|
+
|
|
62
|
+
---
|
|
63
|
+
|
|
64
|
+
## TASK-B (T1): brownfield change in an unfamiliar repo
|
|
65
|
+
|
|
66
|
+
**Shape.** A bug fix or small feature inside a real, mature open-source
|
|
67
|
+
repository that the agent has never been given context for. Not a greenfield
|
|
68
|
+
build. The value being measured is comprehension of code someone else wrote.
|
|
69
|
+
|
|
70
|
+
**Instance selection (mechanical, not curated).** From a public repo pinned at a
|
|
71
|
+
specific commit SHA, select closed issues that (a) have a linked merge commit,
|
|
72
|
+
(b) whose merge commit modifies at least one non-test source file AND at least
|
|
73
|
+
one test file, (c) were closed AFTER the knowledge cutoff of every model under
|
|
74
|
+
test. Sort candidates by issue number and take the first N that satisfy the
|
|
75
|
+
filter. The person tuning the harness does not choose which issues are in the
|
|
76
|
+
set - the filter and the sort order do. See
|
|
77
|
+
[Held-out discipline](#held-out-discipline-mandatory).
|
|
78
|
+
|
|
79
|
+
**What the agent gets.** The repo checked out at the PARENT of the fix commit,
|
|
80
|
+
plus the issue title and body as written by the original reporter. Nothing else.
|
|
81
|
+
The issue text is not rewritten, hinted, or clarified - a rewritten issue is a
|
|
82
|
+
hint, and the whole point of brownfield is that real requests are underspecified.
|
|
83
|
+
|
|
84
|
+
**What is withheld.** The test files added or modified by the fix commit are
|
|
85
|
+
removed from the working tree the agent receives. This is load-bearing and must
|
|
86
|
+
be verified per instance before the run: if the oracle test is present in the
|
|
87
|
+
tree, the agent can read the expected behaviour directly and the oracle measures
|
|
88
|
+
nothing. An instance where the test cannot be cleanly withheld (for example, the
|
|
89
|
+
fix modifies an existing test the agent needs for unrelated context) is
|
|
90
|
+
**dropped from the set and the drop is reported**, not silently patched around.
|
|
91
|
+
|
|
92
|
+
### Oracle
|
|
93
|
+
|
|
94
|
+
Post-hoc, the withheld test files are restored from the fix commit and the
|
|
95
|
+
repo's own test runner is executed against the agent's tree.
|
|
96
|
+
|
|
97
|
+
- **PASS**: every restored test passes AND the repo's pre-existing suite still
|
|
98
|
+
passes at its pinned baseline (no regressions).
|
|
99
|
+
- **PARTIAL**: restored tests pass but the pre-existing suite regresses, or a
|
|
100
|
+
strict subset of restored tests passes.
|
|
101
|
+
- **FAIL**: anything else, including a tree that does not build.
|
|
102
|
+
|
|
103
|
+
This is deterministic, runs offline, costs nothing beyond CPU, and is decided by
|
|
104
|
+
the upstream project's own test code written by the upstream maintainers before
|
|
105
|
+
this benchmark existed.
|
|
106
|
+
|
|
107
|
+
### What this oracle CANNOT catch
|
|
108
|
+
|
|
109
|
+
- **A correct fix that the upstream test does not cover.** Upstream tests encode
|
|
110
|
+
the fix the maintainer actually wrote. A different, equally valid fix can fail
|
|
111
|
+
them. This biases toward the upstream solution and against novel ones.
|
|
112
|
+
- **A fix that passes by coincidence.** Tests can pass on a change that is right
|
|
113
|
+
for the wrong reason and will break on the next input.
|
|
114
|
+
- **Code quality of the diff.** A 400-line change and a 3-line change that both
|
|
115
|
+
pass score identically. Maintainability is not measured here (see NOT
|
|
116
|
+
DETERMINISTIC).
|
|
117
|
+
- **Overfitting to the test's literal assertions** if any part of the test
|
|
118
|
+
leaked into context through a stack trace, CI config, or changelog.
|
|
119
|
+
- **Whether the agent understood the issue** versus pattern-matched a similar
|
|
120
|
+
fix elsewhere in the repo.
|
|
121
|
+
|
|
122
|
+
---
|
|
123
|
+
|
|
124
|
+
## TASK-M (T2): multi-file migration with real coupling
|
|
125
|
+
|
|
126
|
+
**Shape.** Migrate every call site of an internal API to a replacement API
|
|
127
|
+
across a codebase where the call sites differ structurally - some in loops, some
|
|
128
|
+
behind conditionals, some with the old signature's argument order, some already
|
|
129
|
+
partly migrated. Coupling is the point: a task whose files are independent is N
|
|
130
|
+
single-file tasks wearing a trenchcoat.
|
|
131
|
+
|
|
132
|
+
**Instance selection.** Derived from a real migration commit in a pinned public
|
|
133
|
+
repo (same cutoff filter as TASK-B), or from a synthetic-but-fixed corpus
|
|
134
|
+
generated once, frozen, and stored outside this repo. If synthetic, the
|
|
135
|
+
generator's seed is recorded and the corpus is never regenerated between tools -
|
|
136
|
+
a regenerated corpus is a different experiment.
|
|
137
|
+
|
|
138
|
+
**What the agent gets.** The repo at the pre-migration commit, plus the
|
|
139
|
+
migration instruction as a specification of the NEW API's contract (signature,
|
|
140
|
+
semantics, deprecation reason). Not a list of the files to change - finding them
|
|
141
|
+
is the task.
|
|
142
|
+
|
|
143
|
+
### Oracle
|
|
144
|
+
|
|
145
|
+
Three independent deterministic checks, all run post-hoc by the harness. All
|
|
146
|
+
three must hold for PASS.
|
|
147
|
+
|
|
148
|
+
1. **Type checker / compiler.** The project's own `tsc --noEmit`, `mypy`, or
|
|
149
|
+
equivalent at its pinned config. Exit code only. Catches signature
|
|
150
|
+
mismatches and broken call sites.
|
|
151
|
+
2. **Held-out integration tests.** Tests exercising the migrated behaviour
|
|
152
|
+
end-to-end, withheld from the agent's tree exactly as in TASK-B, restored and
|
|
153
|
+
run afterward. These assert BEHAVIOUR is preserved across the migration, not
|
|
154
|
+
just that it compiles.
|
|
155
|
+
3. **Zero-residual-call-site invariant.** An AST query (not a grep - a grep
|
|
156
|
+
matches strings in comments and strings) asserting that zero call sites of
|
|
157
|
+
the old API remain reachable in source.
|
|
158
|
+
|
|
159
|
+
**The invariant must not be satisfiable by deletion.** Check 3 alone is trivially
|
|
160
|
+
passed by deleting every call site. It is only meaningful in conjunction with
|
|
161
|
+
check 2, which fails if the behaviour those call sites provided has disappeared.
|
|
162
|
+
Additionally the harness asserts a **call-site count floor**: the number of NEW
|
|
163
|
+
API call sites must be at least the number of OLD API call sites that existed
|
|
164
|
+
pre-migration. **The consolidation allowance is 0.** An agent may still
|
|
165
|
+
legitimately merge two call sites into one, but each such consolidation is
|
|
166
|
+
reported individually and justified against the behaviour check; a justified
|
|
167
|
+
consolidation does not count against the floor, and an unreported one fails it.
|
|
168
|
+
The allowance is fixed at 0 here rather than left as a threshold to be set
|
|
169
|
+
later, because an unspecified threshold in a pre-registration document is the
|
|
170
|
+
exact defect the frozen-rubric rule below names: a parameter set after seeing
|
|
171
|
+
results is a parameter chosen because of them. Stating this relationship is the
|
|
172
|
+
entire reason check 3 is not reported on its own.
|
|
173
|
+
|
|
174
|
+
- **PASS**: all three checks hold.
|
|
175
|
+
- **PARTIAL**: type check and invariant hold, integration tests partly fail.
|
|
176
|
+
- **FAIL**: type check fails, or call sites remain, or the count floor is
|
|
177
|
+
breached.
|
|
178
|
+
|
|
179
|
+
### What this oracle CANNOT catch
|
|
180
|
+
|
|
181
|
+
- **Semantically wrong migrations that are behaviourally equivalent on the
|
|
182
|
+
tested paths.** Untested paths are unmeasured.
|
|
183
|
+
- **Consolidation that is technically legal but wrong** - the count floor's
|
|
184
|
+
allowance is a judgement call frozen in advance, and a clever wrong answer can
|
|
185
|
+
sit inside it.
|
|
186
|
+
- **Dynamic call sites.** Reflection, string-keyed dispatch, and
|
|
187
|
+
`getattr`-style access are invisible to an AST query. Instances that use them
|
|
188
|
+
should be excluded at selection time, and the exclusion reported.
|
|
189
|
+
- **Migration quality**: whether the resulting code is idiomatic, whether the
|
|
190
|
+
diff is reviewable, whether a human would merge it.
|
|
191
|
+
- **Partial credit for a correct-but-incomplete migration** that stopped at a
|
|
192
|
+
hard file. The oracle is binary per check; PARTIAL is coarse.
|
|
193
|
+
|
|
194
|
+
---
|
|
195
|
+
|
|
196
|
+
## TASK-S (T5): scientific / research task
|
|
197
|
+
|
|
198
|
+
> **THIS ONE IS A PLAN, NOT AN OPERATIONALIZED TASK.**
|
|
199
|
+
>
|
|
200
|
+
> Stated plainly and without hedging: TASK-B and TASK-M can be executed as
|
|
201
|
+
> written by someone who has not read anything else. TASK-S cannot. It requires
|
|
202
|
+
> a specific paper to be chosen, its artifact availability confirmed, and its
|
|
203
|
+
> headline number and tolerance transcribed before it is runnable. What follows
|
|
204
|
+
> is the design and the selection rule, not a runnable instance. Do not report
|
|
205
|
+
> TASK-S results alongside the other two as if they were equally rigorous.
|
|
206
|
+
|
|
207
|
+
**Shape.** Reproduce a published quantitative result. The agent is given a paper
|
|
208
|
+
(or its methods section) and the raw data or the data-acquisition instructions,
|
|
209
|
+
and must produce the reported figure through its own implementation.
|
|
210
|
+
|
|
211
|
+
**Why reproduction rather than open-ended research.** Open-ended research has no
|
|
212
|
+
deterministic oracle - novelty and usefulness require a domain expert, which is
|
|
213
|
+
why 2b lists them as needing a domain judge. Reproduction is the largest subset
|
|
214
|
+
of research work that IS decidable: the answer exists, it is a number, and the
|
|
215
|
+
agent has not been given it.
|
|
216
|
+
|
|
217
|
+
**Instance selection.** A paper with (a) a publicly available dataset, (b) a
|
|
218
|
+
headline numeric result stated with enough precision to compare against, (c)
|
|
219
|
+
publication date after every model's cutoff, or an obscure enough result that
|
|
220
|
+
memorization is implausible - and this second condition is weak, see RISKS. The
|
|
221
|
+
paper's own reference implementation, if one exists, is NOT given to the agent
|
|
222
|
+
and is used only as the differential comparison.
|
|
223
|
+
|
|
224
|
+
**What the agent gets.** The paper's methods section with the results section
|
|
225
|
+
removed, and the dataset. The number it is being scored against is not in its
|
|
226
|
+
context.
|
|
227
|
+
|
|
228
|
+
### Oracle
|
|
229
|
+
|
|
230
|
+
Numeric reproduction within a pre-registered tolerance.
|
|
231
|
+
|
|
232
|
+
- The tolerance is **stated before the run**, derived from the paper's own
|
|
233
|
+
reported variance or confidence interval where one exists, or set at a fixed
|
|
234
|
+
relative tolerance recorded in the frozen rubric where one does not.
|
|
235
|
+
- **PASS**: the agent's produced figure falls within tolerance of the published
|
|
236
|
+
figure.
|
|
237
|
+
- **PARTIAL**: the agent produces a figure by a defensible method that falls
|
|
238
|
+
outside tolerance.
|
|
239
|
+
- **FAIL**: no figure produced, or the pipeline does not run.
|
|
240
|
+
|
|
241
|
+
Secondary deterministic check where a reference implementation exists:
|
|
242
|
+
differential comparison of outputs on identical input. This is a stronger signal
|
|
243
|
+
than the headline number alone, because a single scalar can be hit by accident.
|
|
244
|
+
|
|
245
|
+
### What this oracle CANNOT catch
|
|
246
|
+
|
|
247
|
+
- **Memorization.** If the result is in the training data, reproduction measures
|
|
248
|
+
recall, not method. This is the dominant threat and it is not fully
|
|
249
|
+
mitigable - see RISKS.
|
|
250
|
+
- **Right number, wrong method.** A scalar within tolerance can be reached by a
|
|
251
|
+
pipeline that is wrong in ways that would diverge on any other dataset. The
|
|
252
|
+
differential check reduces this only where a reference implementation exists.
|
|
253
|
+
- **Whether the reproduction is a good scientific artifact** - reusable,
|
|
254
|
+
documented, correct in its statistics.
|
|
255
|
+
- **Novelty or usefulness.** Reproduction measures neither, by construction.
|
|
256
|
+
2b's "Novelty/usefulness" dimension has no oracle here and is not scored.
|
|
257
|
+
- **The paper being wrong.** If the published number is itself an error, the
|
|
258
|
+
oracle rewards reproducing an error.
|
|
259
|
+
|
|
260
|
+
---
|
|
261
|
+
|
|
262
|
+
## Pre-registration
|
|
263
|
+
|
|
264
|
+
Everything in this section is frozen BEFORE the first run. This is the whole
|
|
265
|
+
difference between a benchmark and a story.
|
|
266
|
+
|
|
267
|
+
### QUALITY
|
|
268
|
+
|
|
269
|
+
Scored solely by the per-task oracle above. Three values only: PASS / PARTIAL /
|
|
270
|
+
FAIL, with the per-task definitions as written. No aggregate quality score is
|
|
271
|
+
computed across tasks - three tasks cannot support an average, and a mean over
|
|
272
|
+
three ordinal values is not a measurement.
|
|
273
|
+
|
|
274
|
+
There is **no LLM-judged quality dimension**. If a quality question cannot be
|
|
275
|
+
answered by the oracle, it appears in the NOT DETERMINISTIC list and goes
|
|
276
|
+
unscored.
|
|
277
|
+
|
|
278
|
+
### COST
|
|
279
|
+
|
|
280
|
+
- **Measured**: total USD spend attributable to the run.
|
|
281
|
+
- **Read from**: the per-run receipt's `cost.usd`, surfaced as `cost_usd`, which
|
|
282
|
+
is the convention `tools/cost-per-outcome.py` and `tools/receipt-stats.py`
|
|
283
|
+
already read (`tools/cost-per-outcome.py:151`,
|
|
284
|
+
`tools/receipt-stats.py:138`). This document adopts that convention rather
|
|
285
|
+
than inventing a second one.
|
|
286
|
+
- **UNKNOWN, never 0.** A run that recorded no cost did not cost zero dollars.
|
|
287
|
+
It reads UNKNOWN, it is excluded from every total, and the exclusion is
|
|
288
|
+
reported on the same row as the figure it was excluded from. This rule is
|
|
289
|
+
inherited verbatim from `tools/cost-per-outcome.py`, whose module docstring
|
|
290
|
+
states it as honesty rule 1 and enforces it by letting the MEASURED COUNT
|
|
291
|
+
decide UNKNOWN rather than the total. An unmeasured run is not evidence of
|
|
292
|
+
cheapness; it is an absent measurement.
|
|
293
|
+
|
|
294
|
+
**Cross-tool comparability - the load-bearing caveat.** The above reads Loki's
|
|
295
|
+
artifacts. Competing tools emit no receipt. An oracle that can only read our own
|
|
296
|
+
cost artifact silently advantages us, and a founder reading `$X vs UNKNOWN` will
|
|
297
|
+
read UNKNOWN as "free" or as "broken" when it means neither.
|
|
298
|
+
|
|
299
|
+
The pre-registered rule:
|
|
300
|
+
|
|
301
|
+
1. Cost is captured uniformly where the provider exposes a per-run usage record
|
|
302
|
+
independent of the tool (provider-side billing or usage API, bracketed by the
|
|
303
|
+
harness around the run). That figure, not the tool's self-report, is the
|
|
304
|
+
comparable one.
|
|
305
|
+
2. Where a tool provides no such record, its cost reads **UNKNOWN** for that run.
|
|
306
|
+
3. **A measured figure and an UNKNOWN cannot be differenced, ratioed, or ranked
|
|
307
|
+
against each other.** A cost comparison is reported only across tools that
|
|
308
|
+
were all measured by the same route. If only one tool has a measurable cost,
|
|
309
|
+
there is no cost comparison in that experiment, and the section says so
|
|
310
|
+
instead of showing one column with a number and one column empty.
|
|
311
|
+
|
|
312
|
+
### TIME
|
|
313
|
+
|
|
314
|
+
- **Measured**: wall clock from run start to run end.
|
|
315
|
+
- **Read from**: the tool's own emitted artifacts. For Loki, timestamps in
|
|
316
|
+
`.loki/events.jsonl` (flat `{"timestamp","type","data"}` schema, per
|
|
317
|
+
`autonomy/completion-council.sh:3893`), and where the run produces a preview,
|
|
318
|
+
`.loki/app-runner/first-preview.json`. For tools that emit no timestamped
|
|
319
|
+
artifact, an external timestamp bracket written by the harness immediately
|
|
320
|
+
before spawn and immediately after exit - written by the harness, not by the
|
|
321
|
+
tool.
|
|
322
|
+
- **NEVER `ps etime` on a supervising process.** This is not a stylistic
|
|
323
|
+
preference. This repo hit exactly this bug: a hung benchmark harness reported
|
|
324
|
+
39 minutes for a 20-minute run because `etime` measured the WAITER's lifetime,
|
|
325
|
+
not the work. `pgrep -f` additionally matches the waiter itself. Any time
|
|
326
|
+
figure sourced from a process table is void and is re-measured, not adjusted.
|
|
327
|
+
|
|
328
|
+
**Cross-tool comparability applies to TIME exactly as it does to COST**, and for
|
|
329
|
+
a sharper reason. Our `events.jsonl` clock starts after the engine is already
|
|
330
|
+
up; a harness bracket around a competitor includes process spawn, client init,
|
|
331
|
+
and teardown. Differencing the two makes our number smaller for reasons that are
|
|
332
|
+
not speed - RISK 7 (harness asymmetry) expressing itself inside a scored
|
|
333
|
+
dimension. Therefore: **the comparable wall clock is the harness bracket,
|
|
334
|
+
applied uniformly to every tool INCLUDING ours.** The tool's own artifacts
|
|
335
|
+
(`.loki/events.jsonl`, `.loki/app-runner/first-preview.json`) are the
|
|
336
|
+
finer-grained internal breakdown - reported, never used for a cross-tool
|
|
337
|
+
difference. A harness-written timestamp around spawn and exit is not a
|
|
338
|
+
process-table read, so this satisfies the prohibition above rather than evading
|
|
339
|
+
it.
|
|
340
|
+
|
|
341
|
+
### The frozen rubric
|
|
342
|
+
|
|
343
|
+
The scored dimensions are exactly: **QUALITY** (oracle verdict), **COST**,
|
|
344
|
+
**TIME**. Three. That is the entire rubric.
|
|
345
|
+
|
|
346
|
+
> **A dimension added after seeing results is a dimension chosen because of
|
|
347
|
+
> them.** If, after the first run, someone proposes a fourth dimension, that
|
|
348
|
+
> proposal is evidence about the results and not about the rubric. The honest
|
|
349
|
+
> handling is to run it as a NEW, separately-reported experiment against a
|
|
350
|
+
> re-frozen rubric, and to state in the report that the dimension was added
|
|
351
|
+
> post-hoc and on which observation. Silently widening the rubric converts a
|
|
352
|
+
> benchmark into a search for a framing in which we win.
|
|
353
|
+
|
|
354
|
+
---
|
|
355
|
+
|
|
356
|
+
## Held-out discipline (mandatory)
|
|
357
|
+
|
|
358
|
+
1. **Task instances are not selected by whoever tunes the harness.** Selection
|
|
359
|
+
is by the mechanical filter stated per task - repo pinned at a SHA, filter
|
|
360
|
+
predicate, deterministic sort, first N. The filter is written down before the
|
|
361
|
+
candidate list is inspected. If an instance is dropped (unwithhold-able test,
|
|
362
|
+
dynamic dispatch, missing dataset), the drop and its reason are reported
|
|
363
|
+
alongside the results. An unreported drop is a selection effect.
|
|
364
|
+
|
|
365
|
+
2. **One scored attempt per tool per task.** The first run is the run. A retry
|
|
366
|
+
after seeing the result is a DIFFERENT EXPERIMENT: it is reported separately,
|
|
367
|
+
labelled as a retry, with the observation that prompted it stated. It is
|
|
368
|
+
never averaged into the first attempt, and the first attempt is never
|
|
369
|
+
replaced. "Best of N" where N was chosen after seeing the failures is not a
|
|
370
|
+
measurement of capability; it is a measurement of how many tries the
|
|
371
|
+
experimenter was willing to fund.
|
|
372
|
+
|
|
373
|
+
3. **Task text is NEVER committed to this repo.** A task stored in the
|
|
374
|
+
repository is a task inside the context of the thing being measured - any
|
|
375
|
+
agent running in this tree can read it, and every future model trained on
|
|
376
|
+
this repo has seen it. This document therefore contains task SHAPES,
|
|
377
|
+
SELECTION RULES, and ORACLE DEFINITIONS, and contains no task text, no repo
|
|
378
|
+
URL, no issue number, no paper title, and no expected value.
|
|
379
|
+
|
|
380
|
+
**Where they live instead**: task instances are stored outside this
|
|
381
|
+
repository and outside any directory an agent under test is given - a private
|
|
382
|
+
store, referenced from the run harness by identifier only. The run report
|
|
383
|
+
cites instances by opaque ID. Resolving an ID to its content is a manual step
|
|
384
|
+
performed by the person running the experiment. If an instance's text ever
|
|
385
|
+
appears in this repo, in a commit message, or in a run log committed here,
|
|
386
|
+
that instance is **burned** and is replaced, not reused.
|
|
387
|
+
|
|
388
|
+
---
|
|
389
|
+
|
|
390
|
+
## Calibration caveat (must not be softened)
|
|
391
|
+
|
|
392
|
+
`tools/calibration-audit.py` scores **AGREEMENT WITH THE COUNCIL MAJORITY, NOT
|
|
393
|
+
ACCURACY.** The council's outcome is mechanically derived from the votes
|
|
394
|
+
(`approve_count >= threshold`), so a voter's own prediction partially CAUSES the
|
|
395
|
+
label it is subsequently scored against. A voter scoring perfectly there may
|
|
396
|
+
simply be voting with the crowd. No artifact on disk records whether the council
|
|
397
|
+
was actually RIGHT, so ground-truth calibration is not computable from that
|
|
398
|
+
substrate at all.
|
|
399
|
+
|
|
400
|
+
Therefore, in this document and any report generated from it:
|
|
401
|
+
|
|
402
|
+
- Calibration is **a caveat, never an accuracy term and never a quality term.**
|
|
403
|
+
- Calibration is **never a scored dimension.** It is not in the frozen rubric
|
|
404
|
+
above and may not be added to it. A high calibration figure is not evidence
|
|
405
|
+
that any task was completed well, and must never be cited as though it were.
|
|
406
|
+
|
|
407
|
+
Factual note on availability: `tools/calibration-audit.py` is not present in
|
|
408
|
+
this worktree at the current commit. It exists on another branch (introduced in
|
|
409
|
+
commit `8c33d123`), whose own module docstring states this circularity in the
|
|
410
|
+
same terms. The caveat above stands at full strength regardless of which branch
|
|
411
|
+
the tool is on - it is a prohibition on how a class of number may be used, not a
|
|
412
|
+
description of a file.
|
|
413
|
+
|
|
414
|
+
---
|
|
415
|
+
|
|
416
|
+
## Dimensions that are NOT deterministic and are NOT scored
|
|
417
|
+
|
|
418
|
+
These are inherited from 2b's scored-dimension table, where they are defined.
|
|
419
|
+
They are listed here because this document's job is to say which of them have
|
|
420
|
+
oracles. **These do not.** They are marked NOT DETERMINISTIC and left unscored
|
|
421
|
+
rather than approximated with an LLM judge.
|
|
422
|
+
|
|
423
|
+
| Dimension (defined in 2b) | Why it has no oracle |
|
|
424
|
+
|---|---|
|
|
425
|
+
| Completion quality *beyond the oracle verdict* | 2b defines it as judged against intent. Intent is not a file. The oracle covers "does it pass"; "is it what was wanted" is human. |
|
|
426
|
+
| Maintainability | Requires reading the diff as a future maintainer. No checker distinguishes a clean fix from a working mess. |
|
|
427
|
+
| Intervention rate | Countable in principle, but what counts as an intervention is a judgement about whether a nudge was necessary or merely habitual. |
|
|
428
|
+
| Novelty / usefulness (T5) | 2b already states this needs a domain judge. Reproduction deliberately sidesteps it and therefore does not measure it. |
|
|
429
|
+
| User value ("would the requester ship it") | 2b's own framing: the only dimension that ends in a human. It ends in a human here too. |
|
|
430
|
+
| Recovery after induced failure | Requires deliberate fault injection not designed here; the fault set would be chosen by the harness tuner, which is the selection effect this document exists to avoid. |
|
|
431
|
+
| Verification tax | The time/token half is measurable; "how often it changed an outcome" requires knowing the counterfactual outcome, which is not observable from a single run per task. |
|
|
432
|
+
|
|
433
|
+
Being unscored is not the same as being unimportant. Several of these matter
|
|
434
|
+
more than the ones that are scored. They are excluded because a fabricated
|
|
435
|
+
measurement of an important thing is worse than an acknowledged gap.
|
|
436
|
+
|
|
437
|
+
---
|
|
438
|
+
|
|
439
|
+
## RISKS: what could make these results misleading even if executed perfectly
|
|
440
|
+
|
|
441
|
+
1. **Training-data contamination.** The dominant threat. Public repos, their
|
|
442
|
+
issues, their fixes, and published papers with their numbers are all
|
|
443
|
+
plausibly in training data. Post-cutoff filtering is a mitigation, not a
|
|
444
|
+
solution: cutoffs are approximate, self-reported, and differ per tool, so a
|
|
445
|
+
date filter that holds one model out may not hold another. A tool that
|
|
446
|
+
memorized the fix and a tool that derived it produce identical oracle
|
|
447
|
+
verdicts. **This affects TASK-S most severely** and is why TASK-S results, if
|
|
448
|
+
ever produced, carry lower evidential weight than TASK-B or TASK-M.
|
|
449
|
+
|
|
450
|
+
2. **n is far too small for a ranking.** Three task shapes, one instance each,
|
|
451
|
+
one attempt each. This can support existence claims ("tool X solved this
|
|
452
|
+
instance") and cannot support comparative claims ("tool X is better than tool
|
|
453
|
+
Y"). Any percentage computed over three tasks is theatre. Reporting a winner
|
|
454
|
+
from this design would be the central misuse.
|
|
455
|
+
|
|
456
|
+
3. **Selection effects that survive the mechanical filter.** The filter is
|
|
457
|
+
mechanical but the CHOICE OF FILTER is not - which repo, which language,
|
|
458
|
+
which issue-shape predicate. A filter that happens to select tasks resembling
|
|
459
|
+
our own test corpus advantages us. Mitigation: fix the repo and filter before
|
|
460
|
+
inspecting candidates, and report the drop list.
|
|
461
|
+
|
|
462
|
+
4. **Oracle gaming.** A held-out test suite is a target once the agent infers it
|
|
463
|
+
exists. Agents can write code that special-cases plausible test inputs, or
|
|
464
|
+
over-fit to behaviour implied by the issue text. TASK-M's residual-call-site
|
|
465
|
+
invariant is specifically gameable by deletion, which is why it is never
|
|
466
|
+
reported without the behaviour check and the count floor. Assume any oracle
|
|
467
|
+
stated publicly will eventually be optimized against.
|
|
468
|
+
|
|
469
|
+
5. **Tool-version drift mid-experiment.** These CLIs ship frequently - this repo
|
|
470
|
+
itself releases hourly. A run of tool A on Monday and tool B on Wednesday
|
|
471
|
+
compares two different weeks, not two tools. Mitigation: pin and record the
|
|
472
|
+
exact version of every tool, and re-run the whole matrix if any version
|
|
473
|
+
changes mid-experiment. A partially-refreshed matrix is void.
|
|
474
|
+
|
|
475
|
+
6. **Environment non-determinism.** Network flakiness, rate limits, provider-side
|
|
476
|
+
model routing, and nondeterministic sampling mean the same tool on the same
|
|
477
|
+
task can produce different outcomes. With one attempt per task, a single rate
|
|
478
|
+
limit is indistinguishable from a capability failure. Rate-limit and
|
|
479
|
+
infrastructure failures must be recorded as INFRASTRUCTURE, not FAIL, and
|
|
480
|
+
such a run is re-run - which is itself a deviation from one-attempt
|
|
481
|
+
discipline and must be reported as one.
|
|
482
|
+
|
|
483
|
+
7. **Harness asymmetry.** We wrote the harness. Our tool's artifacts are
|
|
484
|
+
first-class to it; competitors are driven through whatever CLI surface they
|
|
485
|
+
expose. Effort spent making our own path work smoothly and not theirs is an
|
|
486
|
+
advantage that looks like capability. Mitigation: drive every tool through
|
|
487
|
+
its documented non-interactive entry point only, and report any tool for
|
|
488
|
+
which the harness needed tool-specific accommodation.
|
|
489
|
+
|
|
490
|
+
8. **The oracle can be wrong.** Upstream tests can be flaky, papers can report
|
|
491
|
+
errors, and type checkers have config-dependent behaviour. A deterministic
|
|
492
|
+
oracle is reproducible, which is not the same as correct.
|
|
493
|
+
|
|
494
|
+
---
|
|
495
|
+
|
|
496
|
+
## COST-TO-RUN estimate
|
|
497
|
+
|
|
498
|
+
**This is an estimate, not a measurement.** No run has occurred; nothing here is
|
|
499
|
+
read from an artifact. It exists so the expenditure can be approved or declined
|
|
500
|
+
before rather than after.
|
|
501
|
+
|
|
502
|
+
**Assumptions** (each is a place the estimate can be wrong):
|
|
503
|
+
|
|
504
|
+
- 3 tasks x 1 attempt x the number of tools compared. At 4 tools that is 12
|
|
505
|
+
scored runs.
|
|
506
|
+
- Per-run agentic spend of roughly $2-$15 depending on tool, model tier, and
|
|
507
|
+
iteration count. The spread is wide because iteration count is the dominant
|
|
508
|
+
cost term and is exactly what the experiment is measuring - it cannot be
|
|
509
|
+
known in advance without circularity.
|
|
510
|
+
- Oracle execution (test runs, type checks, AST queries) is local CPU only:
|
|
511
|
+
effectively $0 in provider spend.
|
|
512
|
+
- TASK-S is excluded from this estimate because it is a plan, not a runnable
|
|
513
|
+
task. Adding it means adding dataset acquisition and compute of unknown size.
|
|
514
|
+
- Setup labour (instance selection, test-withholding verification, harness
|
|
515
|
+
wiring per tool) is human time and is NOT included in the dollar figure.
|
|
516
|
+
|
|
517
|
+
**Estimated provider spend**: roughly **$25 to $180** for a 4-tool, 2-task
|
|
518
|
+
(TASK-B and TASK-M) matrix. The upper bound is what to budget; the lower bound
|
|
519
|
+
is what to hope for.
|
|
520
|
+
|
|
521
|
+
**Estimated wall clock**: roughly **4 to 12 hours** of run time for the matrix
|
|
522
|
+
if run serially, plus an estimated **1 to 3 days of human setup** dominated by
|
|
523
|
+
instance selection and by verifying per instance that the oracle tests are
|
|
524
|
+
genuinely withheld. The setup is the expensive half and is not parallelizable by
|
|
525
|
+
adding tools.
|
|
526
|
+
|
|
527
|
+
**What the money buys.** Existence claims on two task shapes with
|
|
528
|
+
non-self-graded oracles, and a reusable harness. It does NOT buy a ranking - see
|
|
529
|
+
RISKS 2. A founder approving this should approve it as instrumentation, not as a
|
|
530
|
+
competitive result, and should decline it if a ranking is what is wanted, because
|
|
531
|
+
this design cannot produce one honestly.
|
|
532
|
+
|
|
533
|
+
---
|
|
534
|
+
|
|
535
|
+
*Status per section 2b's Status separation table: **Planned experiments**. Zero
|
|
536
|
+
runs to date.*
|
|
@@ -0,0 +1,97 @@
|
|
|
1
|
+
# Prompt ablation: measured result
|
|
2
|
+
|
|
3
|
+
Date: 2026-08-03. Instrument: `benchmarks/run-prompt-ablation.sh`,
|
|
4
|
+
reporter `benchmarks/report-prompt-ablation.py`. Raw rows:
|
|
5
|
+
`benchmarks/results/prompt-ablation.jsonl`.
|
|
6
|
+
|
|
7
|
+
## The question
|
|
8
|
+
|
|
9
|
+
Anthropic deleted roughly 80% of Claude Code's system prompt for Opus 5, on
|
|
10
|
+
the finding that instructions written to correct older models had become dead
|
|
11
|
+
weight, and that the model measured slightly MORE capable without them.
|
|
12
|
+
|
|
13
|
+
`LOKI_SIMPLE=1` strips the coaching half of our prompt: the RARV cycle, SDLC
|
|
14
|
+
phases, memory habits. Per-iteration state (which gate failed, self-heal
|
|
15
|
+
output, checklist status) is never touched, because that is information the
|
|
16
|
+
model cannot derive from anywhere else.
|
|
17
|
+
|
|
18
|
+
Prompt size fell 78% (about 1562 tokens per iteration), reproduced
|
|
19
|
+
independently on both the bash and Bun routes. That is a TOKEN measurement.
|
|
20
|
+
It says nothing about whether the stripped arm builds the same thing, faster
|
|
21
|
+
or slower, more or less reliably. This is that measurement.
|
|
22
|
+
|
|
23
|
+
## Result
|
|
24
|
+
|
|
25
|
+
```
|
|
26
|
+
trials recorded: full=3 simple=3
|
|
27
|
+
|
|
28
|
+
metric full (default) simple (LOKI_SIMPLE=1)
|
|
29
|
+
-----------------------------------------------------------
|
|
30
|
+
wall clock 8.2 min [7.3-8.7] n=3 7.6 min [6.2-8.0] n=3
|
|
31
|
+
iterations 1.0 [1.0-1.0] n=3 1.0 [1.0-1.0] n=3
|
|
32
|
+
|
|
33
|
+
full reliability: 3/3 completed
|
|
34
|
+
simple reliability: 3/3 completed
|
|
35
|
+
|
|
36
|
+
VERDICT: NO DIFFERENCE DEMONSTRATED -- the arms' ranges overlap
|
|
37
|
+
(full 7.3-8.7, simple 6.2-8.0), so the median gap is within
|
|
38
|
+
the run-to-run noise.
|
|
39
|
+
```
|
|
40
|
+
|
|
41
|
+
## What this means, stated carefully
|
|
42
|
+
|
|
43
|
+
**Speed: no difference demonstrated.** The simple arm's median is 0.6 min
|
|
44
|
+
lower, but the ranges overlap heavily. A gap smaller than the variance each
|
|
45
|
+
arm shows against itself is not a finding.
|
|
46
|
+
|
|
47
|
+
**Reliability: no difference demonstrated.** 3/3 completed in both arms, and
|
|
48
|
+
every run passed its acceptance check. Both arms finished in a single
|
|
49
|
+
iteration.
|
|
50
|
+
|
|
51
|
+
**Cost: the token saving is real and stands on its own.** 78% less prompt per
|
|
52
|
+
iteration is a deterministic, reproducible measurement that does not depend on
|
|
53
|
+
these trials at all.
|
|
54
|
+
|
|
55
|
+
**The honest summary: stripping the coaching prompt cost us nothing
|
|
56
|
+
measurable, and saved 78% of the prompt.** That is a weaker claim than "the
|
|
57
|
+
model is better without it" and a stronger one than "no effect" -- removing
|
|
58
|
+
1562 tokens of instruction per iteration did not degrade speed, completion, or
|
|
59
|
+
acceptance across six real builds.
|
|
60
|
+
|
|
61
|
+
## Why this was nearly reported wrong
|
|
62
|
+
|
|
63
|
+
After trial 1, the numbers were full 8.2 min against simple 6.2 min. That is a
|
|
64
|
+
24% improvement, and it is exactly the shape of result that gets written into
|
|
65
|
+
a release note.
|
|
66
|
+
|
|
67
|
+
Trial 2 reversed it: full 7.3, simple 8.0.
|
|
68
|
+
|
|
69
|
+
One trial of a stochastic agent is indistinguishable from noise. The reporter
|
|
70
|
+
refuses to declare a winner when the ranges overlap, and refuses entirely
|
|
71
|
+
below n=3 per arm, which is why the instrument was built before the claim
|
|
72
|
+
rather than after it.
|
|
73
|
+
|
|
74
|
+
## Limits, stated rather than left for a reader to discover
|
|
75
|
+
|
|
76
|
+
- **n=3 per arm.** Enough to see that the arms overlap; not enough to detect a
|
|
77
|
+
small real effect. A difference under roughly 20% would not be visible here.
|
|
78
|
+
- **One spec, one model.** A single small build on sonnet. Coaching may matter
|
|
79
|
+
more on a longer task, a weaker model, or a spec where the RARV structure is
|
|
80
|
+
load-bearing. This says nothing about those.
|
|
81
|
+
- **Both arms finished in one iteration**, so this never exercised the
|
|
82
|
+
iterate-and-recover path where per-iteration state matters most. That is the
|
|
83
|
+
case where stripping coaching is most likely to be safe and where the
|
|
84
|
+
remaining prompt does the most work, and it is untested.
|
|
85
|
+
- **Degraded providers are unaffected.** Codex and Aider take an earlier
|
|
86
|
+
return path, so the flag never reaches its gate there. Measured as a 0-byte
|
|
87
|
+
delta, not assumed.
|
|
88
|
+
|
|
89
|
+
## Reproduce
|
|
90
|
+
|
|
91
|
+
```bash
|
|
92
|
+
benchmarks/run-prompt-ablation.sh --trials 3 --model sonnet --max-iters 6
|
|
93
|
+
python3 benchmarks/report-prompt-ablation.py benchmarks/results/prompt-ablation.jsonl
|
|
94
|
+
```
|
|
95
|
+
|
|
96
|
+
Both arms run in ONE engine copy; the only difference is the environment
|
|
97
|
+
variable read at prompt-assembly time.
|