dreamteamer 0.20.0 → 0.22.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,435 @@
1
+ # proofs — what a skill, a command or a script CLAIMS, and whether it still holds
2
+
3
+ A **proof** is a record that says what an artifact does, in a form the engine can run and judge.
4
+ `dt prove <id>` runs it and answers with an **exit code**, so a hook, a CI step or a session can
5
+ branch on the result without reading prose.
6
+
7
+ Two things make it different from a test suite. It is about an **artifact** — a skill, a command,
8
+ a command-binding, a module script — not about a function; and its subject is usually a
9
+ **procedure a session performs**, which no test runner can call. So a proof may stop, ask for the
10
+ action to be taken, and be re-run to judge the state that action left behind.
11
+
12
+ | the question | read |
13
+ |---|---|
14
+ | gate, live, or neither | the three kinds |
15
+ | the file and its keys | authoring a proof |
16
+ | what `where`, `count` and `{record}` mean | the predicate language |
17
+ | it stopped and printed PERFORM | the protocol |
18
+ | a proof that WRITES records | the sandbox |
19
+ | where the evidence lives | the ledger |
20
+ | what the number means | exit codes |
21
+ | coverage — what has no proof | reading the surface |
22
+ | proving a skill actually TEACHES | the eval layer |
23
+
24
+ ## the three kinds
25
+
26
+ **A gate** is a static check with no record and no post-state: one or more `run` steps, judged on
27
+ their exit codes alone. `npm test`, a linter, a script's `--dry-run`, `dt check`. It is the cheapest
28
+ proof there is and the only one that needs nothing from the workspace's data. A gate takes no
29
+ `mode`, no `given` and no `expect` — compile refuses all three, because a gate that carried them
30
+ would be a live proof whose author believes something untrue about what runs.
31
+
32
+ **A live proof** runs against a real record and is judged by `expect`. It picks ONE record (`given`),
33
+ snapshots what it is about to measure, runs its steps, and then asks the store whether the world
34
+ changed the way the proof says it should. `mode: readonly` reads the workspace; `mode: writes`
35
+ mutates records and therefore runs in a throwaway worktree (see the sandbox). A live proof is where
36
+ a `perform` step belongs: "run `/close-note` on this note" is an instruction to an actor, and the
37
+ engine's job is to hold the before-state, hand over the instruction, and judge the after-state.
38
+
39
+ **An eval** is the third thing, and the engine deliberately does NOT run it. Whether a skill actually
40
+ teaches — whether a fresh session finds it, loads it and does the job right — is answered by running
41
+ sessions and scoring them, not by a filter over records. It is a PROCEDURE (below), with a scoring
42
+ sheet, run by a human or an agent. There is no `kind: eval`; `PROOF_KINDS` is `gate` and `live`, and
43
+ a workspace that pretended otherwise would be claiming an automated answer to a question nothing
44
+ automated can ask.
45
+
46
+ ## authoring a proof
47
+
48
+ `modules/<module>/proofs/<id>.proof.yaml`. The filename is the id and must equal `name`. ⚠ **`dt add
49
+ proofs` is refused on purpose** — a proof is hand-authored, like a skill or a command, and the
50
+ refusal prints the path to write. Compile validates every key it interprets and FAILS on a proof it
51
+ cannot interpret; run `dt compile` after writing one.
52
+
53
+ ```yaml
54
+ name: notes-close-cleanly # required — equals the filename
55
+ about: [commands/close-note] # required, ≥1 — skills/<id> · commands/<id> ·
56
+ # command-bindings/<id> · <module-id>/bin/<file>
57
+ kind: live # required — gate | live
58
+ mode: readonly # required on live, forbidden on gate — readonly | writes
59
+ description: closing a note sets its status and leaves the body alone.
60
+ external: false # optional — true excludes it from a bare --all
61
+ requires: { env: [HR_EXPORT_DIR], bin: [jq] } # optional
62
+ given: # live only — exactly ONE of `where` or `fixture`
63
+ collection: notes
64
+ where: { status: { _eq: open } }
65
+ pick: latest # `latest` (the collection's sort_field, DESC) or an explicit id
66
+ steps: # required — gate: ≥1 `run`; live: ≥1 of either
67
+ - perform: /close-note {record}
68
+ expect: # live only, ≥1
69
+ - { record: '{record}', where: { status: { _eq: done } } }
70
+ timeout: 120 # optional — seconds per `run` step (default 120)
71
+ ```
72
+
73
+ A gate is smaller, and most proofs in a healthy workspace are gates:
74
+
75
+ ```yaml
76
+ name: hr-export-runs-clean
77
+ about: [hr/bin/export.mjs]
78
+ kind: gate
79
+ description: the export script runs end to end on the sample input and exits 0.
80
+ requires: { bin: [node] }
81
+ steps:
82
+ - run: node modules/hr/bin/export.mjs --dry-run
83
+ - run: node modules/hr/bin/export.mjs --check
84
+ ```
85
+
86
+ A `writes` proof brings its own records, under `modules/<module>/proofs/fixtures/<proof-id>/`,
87
+ mirroring the workspace root — so `data/notes/fx-open.note.md` under that folder becomes
88
+ `notes/fx-open` inside the sandbox:
89
+
90
+ ```yaml
91
+ name: closing-a-note-writes-the-field
92
+ about: [commands/close-note]
93
+ kind: live
94
+ mode: writes
95
+ given: { collection: notes, fixture: true, pick: fx-open }
96
+ steps:
97
+ - perform: /close-note {record}
98
+ expect:
99
+ - { record: '{record}', where: { status: { _eq: done } } }
100
+ - { collection: notes, where: { status: { _eq: done } }, count: { _delta: 1 } }
101
+ ```
102
+
103
+ | key | notes |
104
+ |---|---|
105
+ | `about` | what this proof is FOR. An unresolvable ref FAILS compile — a proof about nothing reports success forever |
106
+ | `kind` | `gate` (run steps only) or `live` (a record and expectations) |
107
+ | `mode` | live only. `writes` implies a sandbox unless `--here` |
108
+ | `requires` | `env` names must be declared in `dreamteamer.vars` or a module's `dreamteamer.env`; `bin` is looked up on PATH. Unmet ⇒ **UNAVAILABLE**, never FAIL — this machine cannot answer the question, which is not a fact about the artifact. ⚠ **names only** — no `.env` VALUE is ever read, compared or printed |
109
+ | `given` | exactly one of `where` (a live filter) or `fixture: true` (this proof's own records). `pick: latest` needs a `sort_field` on the collection; `pick: <id>` names one; **`pick: any` is refused** — a proof that picks arbitrarily proves something arbitrary. A fixture needs `pick: <id>` |
110
+ | `steps` | in order, until one fails or one asks for an actor |
111
+ | `expect` | live only. Four forms, below |
112
+ | `external` | for a proof needing the network, a credential or a mount. Invisible to `--all`; `--external` includes it |
113
+ | `timeout` | positive integer seconds, per `run` step |
114
+
115
+ ## the predicate language
116
+
117
+ **`where` is the ordinary filter grammar** — the same one `dt list --where`, ui-views and binding
118
+ gates use, enumerated in `dt help` and described in `records.md`. Two proof-specific rules:
119
+
120
+ - **one hop, and a second is refused.** `{ owner: { name: { _eq: Ada } } }` resolves `owner` as a
121
+ reference and tests the target's `name`. A third level is refused at compile —
122
+ `where hops more than one reference (<field>.<hop>.<key>) — a proof filter hops at most one` —
123
+ because the evaluator would silently narrow it to false. A proof that genuinely needs two hops
124
+ writes its `expect` against the far collection.
125
+ - ⚠ **a nested key is a reference HOP, not a field comparison.** `{ a: { b: … } }` never means
126
+ "compare `a` to `b`" — it means "resolve `a`, then test `b` on the target". When `a` is not a
127
+ reference field, compile says so by name; at run time it would narrow to zero rows with no
128
+ warning at all, which is the silent failure this whole kind exists to remove.
129
+
130
+ Compile also checks every literal against a CLOSED enum, so `status: { _eq: closed }` on a
131
+ `[open, done]` field is a compile error rather than a filter that matches nothing forever.
132
+
133
+ **Substitution.** Only two braces are substituted, in `run` and `perform` strings: `{record}` becomes
134
+ the picked record's reference (`notes/b`), and `{record.<field>}` becomes that field's value. ⚠
135
+ **Everything else reaches the shell as written** — `awk '{print $1}'`, `sed -n '1,3{p}'`, `jq '{a: .b}'`
136
+ and `mkdir -p x/{a,b}` are all correct steps, and refusing them was measured to be worse than the
137
+ typo the refusal was hunting. So the net is advisory: **compile WARNS** on an identifier-shaped brace
138
+ nobody substitutes (`⚠ proofs/x.proof.yaml: step 2 uses "{recrod}" — only {record} and {record.<field>} are
139
+ substituted; the rest reaches the shell as written`)
140
+ and the proof still compiles and still runs. `${…}` is untouched — that bracket is the resolver's
141
+ (`${env:FILES_FOLDER}`) and the shell's (`${HOME}`).
142
+
143
+ ⚠ **A `path:` and every string literal inside an expectation's `where` are STRICT.** There the
144
+ ENGINE consumes the string, so an unknown brace THROWS instead of passing through:
145
+ `path: "{recrod}/out.txt"` would otherwise become a literal directory that does not exist and the
146
+ expectation would answer `exists false` — a FAIL naming the wrong cause; a filter literal fails the
147
+ same way one layer quieter, becoming a value the field never equals.
148
+
149
+ ⚠ **A `where`'s literals are substituted BEFORE the filter runs, and that is what makes the
150
+ commonest live proof work at all.** `where: { owner: { _eq: "{record}" } }` counts the records
151
+ pointing back at the one this proof picked — the shape most collection-scope expectations take. It
152
+ is rendered in all three places a filter is evaluated: the `_delta` snapshot, the pre-check, and the
153
+ after-pass. ⚠ Inside a `_delta` write **`{record}` (the reference), not `{record.<field>}`**: the two
154
+ counts are taken either side of the steps, so a field literal is rendered from two different values
155
+ and their difference means nothing — assert a field with a `record:` expectation instead.
156
+
157
+ ⚠ **`given.where` may NOT use `{record}`** — compile refuses it (`given.where cannot use {record} —
158
+ the given is what PICKS the record`). The substitution is an EXPECTATION's, and for a reason that is
159
+ not a rule but arithmetic: the given is what selects the record, so at the moment its filter runs
160
+ there is nothing bound yet. The literal used to compile and the proof then answered NO-FIXTURE
161
+ forever, which reads as a fact about the workspace's data. For the same reason, a `{record}` literal
162
+ in an expectation's `where` on a proof with **no `given` at all** is refused too (`an expectation
163
+ uses {record} but this proof declares no given`).
164
+
165
+ ⚠ **`record:` is not a template — it is always the literal `{record}`**, and compile refuses any
166
+ other value (`a record expectation targets {record} — the picked record is its only target`). The
167
+ `given` picks one record and there is no second one to target, so `record: notes/b` was a proof
168
+ judged against a different record than the one it names.
169
+
170
+ **The four expectation forms.**
171
+
172
+ | form | asserts |
173
+ |---|---|
174
+ | `{ collection, where, count }` | how many records of `collection` match `where`, collection-scope |
175
+ | `{ record: '{record}', where }` | the picked record's own fields. Needs a `given`; an empty `where` is refused |
176
+ | `{ step: <n>, exit, stdout, stdout_json }` | what one `run` step did. `step` defaults to the LAST run step; an index past the last step is refused |
177
+ | `{ path: '<template>', exists: true\|false }` | whether a path is there, rendered through the ONE resolver, **relative to the workspace root** (the sandbox's root inside a `writes` proof) |
178
+
179
+ ⚠ **A row is exactly ONE form, and a row whose keys span two is a compile error** —
180
+ `expect[<i>] mixes two forms — a row is one of collection+where+count · record+where · step ·
181
+ path+exists`. It is refused rather than resolved because compile and the judge would otherwise read
182
+ the same row as different shapes, and a row judged as the wrong form produces **zero verdict lines**
183
+ — which passes, vacuously. Measured: `{ record: '{record}', path: 'nope.txt', exists: true }`
184
+ compiled as a path row, was judged as a record row, and answered exit 0 `PASS` on a file that has
185
+ never existed. `{ record, where, count }` was the same seam upside down — the count's operators and
186
+ integers were validated line by line and then never read.
187
+
188
+ **`count`** takes its own closed operator set — `_eq _neq _gt _gte _lt _lte _delta` — and every
189
+ operand must be an INTEGER: `_gte: 'one'` and `_eq: 1.5` are filters that can never be satisfied, and
190
+ compile says so as `count "<op>" compares "<v>", which is not an integer`. A bare scalar is the
191
+ `_eq` it stands for. `count: {}` is refused: it compares nothing and so holds for every possible
192
+ count.
193
+
194
+ **`_delta` is after − before**, judged against the snapshot taken before the first step ran and
195
+ carried on the PENDING row. It is the honest way to say "one more note exists" in a collection that
196
+ already has records — `count: { _gte: 1 }` there is true before anything happens. ⚠ A missing
197
+ snapshot is a **FAIL** naming the repair (`no before-count in the pending row — re-run dt prove <id>
198
+ --restart`), never a delta measured from zero: fail-open there turned "one more" into "at least one"
199
+ and passed on a collection nothing had touched.
200
+
201
+ **`stdout` and `stdout_json`.** A step's stdout is captured up to **64 KB** and judged WHOLE, not as
202
+ a tail. `stdout` takes filter operators over the text (`{ _contains: 'wrote 3 rows' }`);
203
+ `stdout_json` parses the full capture and takes a dotted path per condition
204
+ (`{ 'summary.rows': { _gte: 1 } }`). A payload that is not JSON is its own verdict —
205
+ `stdout is not JSON (…) ✖` — rather than a pile of `undefined` comparisons. Output larger than
206
+ 64 KB belongs in a file the proof asserts with `path:`.
207
+ ⚠ **Both are FILTER MAPS, and compile refuses anything else.** `stdout: hello` reads like "the step
208
+ printed hello" and is the spelling everybody tries first — it used to compile clean, produce zero
209
+ verdict lines, and PASS against a step that printed something else. Write `{ _contains: 'hello' }`.
210
+
211
+ **Every verdict line prints the ACTUAL value beside the wanted one**, always:
212
+
213
+ ```
214
+ count +1 = +1 ✔
215
+ status "open" ∈ [done] ✖
216
+ exit 0 = 0 ✔
217
+ ```
218
+
219
+ ## the protocol — run it, do what it says, run it again
220
+
221
+ Most live proofs about a command or a skill cannot be run by a machine end to end: the middle step
222
+ is an actor. So the runner stops there, records a PENDING row, and prints the block:
223
+
224
+ ```
225
+ in /repo/.worktrees/.tmp-a1b2c3 (a throwaway worktree — records written here are never landed)
226
+ PERFORM /close-note notes/b
227
+ source modules/default/commands/close-note.command.md
228
+ then dt prove notes-close-cleanly --record notes/b (the same verb, again)
229
+ ```
230
+
231
+ - **`in`** appears only when there IS a sandbox, and it comes FIRST — "where am I acting" has to be
232
+ read before "what do I do", or a human writes real records the proof will never judge.
233
+ - **`PERFORM`** is the instruction, substituted. Take it.
234
+ - **`source`** is the command's SOURCE file, off the manifest, when the text opens with `/<command-id>` —
235
+ so the actor can open the thing they are being asked to run.
236
+ - **`then`** is the exact line to type next. The same verb, again: `--record` names the pending run
237
+ this invocation is FINISHING, and nothing else.
238
+
239
+ The re-run judges — it does not re-run the steps. The record, the step results and the `_delta`
240
+ snapshot all come off the pending row, because those are the facts of the run being finished.
241
+
242
+ **A live pending run blocks a fresh one.** Any of them: the refusal names the record to finish
243
+ with, and when several records are pending it lists them all rather than sending you round the loop
244
+ once per row. (The one case that resumes itself is a proof with no `given` — it pends against no
245
+ record, so `--record` cannot name it and the bare verb is the only way back.)
246
+
247
+ **`--restart`** discards every live pending run and starts over: the discarded rows are recorded as
248
+ discarded and their sandboxes removed. ⚠ A discarded row is written as a **FAIL**, so until the
249
+ restarted run reaches a verdict this proof's ledger tail is a FAIL — `dt status --strict` is red in
250
+ between, which is correct (nothing has passed since) and worth knowing before you wire it into a
251
+ hook.
252
+
253
+ A pending run older than the proof's own `timeout` is **stale** — the process it belonged to is gone
254
+ — so it is cleared out loud on the next run.
255
+
256
+ ## the sandbox — where a `writes` proof is allowed to write
257
+
258
+ `mode: writes` mutates records, so it never touches the invoking store. The runner cuts a detached
259
+ throwaway worktree (`.worktrees/.tmp-*`), copies the proof's fixture into it, runs there, judges
260
+ there, and removes it.
261
+
262
+ - **A sandbox is cut from HEAD.** Uncommitted sources are ABSENT from it — the engine-ships-first
263
+ rule in miniature. A collection whose descriptor is not committed is not compiled inside the
264
+ sandbox, its records are invisible rather than invalid, and every count would read zero. So the
265
+ runner refuses first, naming it: `collection "ghosts" is not compiled in the sandbox — commit its
266
+ descriptor, because a sandbox is cut from HEAD`.
267
+ - **A fixture may contain only `data/`.** It mirrors the workspace ROOT, so anything else would
268
+ overwrite what the checkout carries — including the very descriptors its records are then
269
+ validated against. Dot-entries are ignored (your file manager writes `.DS_Store` there and it comes
270
+ back), and `data` must be a directory.
271
+ - **The fixture is the one input nothing validated on the way in**, so the engine's own `check` runs
272
+ against the sandbox before any step does; the first violation is what the failure names.
273
+ - **`--keep`** leaves the sandbox in place after a verdict, to look at. It is only meaningful for a
274
+ sandboxed proof, and `dt status` counts what runs left behind — `.worktrees/` is gitignored, so
275
+ the ledger is the only thing that knows the directory exists. A removal that FAILED is recorded as
276
+ `sandbox_removed: false` and counted the same way.
277
+ - **`--here` runs a `writes` proof in THIS checkout**, and says so out loud before anything moves. It
278
+ is legitimate for exactly one shape: a `given.where` writes proof, against a real record, when the
279
+ point is to prove the thing on live data. A `writes` proof with no fixture runs ONLY that way —
280
+ otherwise it is refused at run time, naming both ways forward.
281
+
282
+ ## the ledger — per machine, per proof, and disposable
283
+
284
+ `.dreamteamer/.proofs/<proof-id>.jsonl` — one JSON row per run, oldest first, appended, **capped at
285
+ the last 50**. It answers "when did this last pass, on this machine, and against which record", and
286
+ it is what makes a `perform` step resumable.
287
+
288
+ ⚠ **The dot is load-bearing.** `.dreamteamer/proofs/` is the compiled KIND folder and compile wipes
289
+ every kind folder on every run, so a ledger written there would be destroyed silently. The
290
+ dot-prefixed sibling is invisible to that loop and to every source enumeration, and it is already
291
+ gitignored.
292
+
293
+ **It is evidence, not data.** `rm -rf .dreamteamer && dt compile` — the folk recovery for a stale
294
+ runtime — takes the ledger with it, and the cost is **re-proving, not data loss**: what a proof
295
+ asserts lives in the committed source, and the rows were only ever about this machine. A malformed
296
+ line (a killed run, a hand edit) is skipped with a warning and repaired by the next append.
297
+
298
+ ## exit codes
299
+
300
+ The point of the verb: a script branches on the number.
301
+
302
+ | code | state | means |
303
+ |---|---|---|
304
+ | `0` | PASS | every expectation held |
305
+ | `1` | FAIL | a step failed, or an expectation did not hold |
306
+ | `2` | usage | RESERVED, and not reachable from a proof — it is what a RETIRED verb spelling answers. A bad flag or an unknown target inside `dt prove` is an ordinary error at `1` |
307
+ | `3` | UNAVAILABLE | this machine lacks a required var or binary — **not** a failure of the artifact |
308
+ | `4` | NO-FIXTURE | the `given` matched no record, or the fixture folder holds none |
309
+ | `5` | PENDING | a `perform` step is owed a human or an agent |
310
+ | `6` | VACUOUS | every expectation ALREADY held before any step ran — or the run produced NO verdict at all |
311
+
312
+ **VACUOUS is the most valuable state in the set.** A proof whose expectations already hold reports
313
+ PASS forever and measures nothing — the silent green this whole verb exists to remove. It is checked
314
+ BEFORE any step runs, so it is caught before the proof takes an action. `step` and `path`
315
+ expectations are not pre-checkable and never make a proof vacuous.
316
+
317
+ ⚠ **And the floor under it: a run that DECLARED expectations and produced zero verdict lines is
318
+ VACUOUS too** — `no expectation produced a verdict — a proof that asserts nothing is not a proof`.
319
+ `verdicts.every(ok)` is vacuously true over an empty list, so every shape that made a declared
320
+ expectation judge nothing came out as `PASS` at exit 0. Compile refuses those shapes now; the floor
321
+ is what makes the class unreachable rather than closed one spelling at a time, including for a
322
+ runtime an older engine wrote. A `gate` declares no expectations at all — its assertion is the
323
+ step's exit code — and is untouched.
324
+
325
+ ⚠ **`dt prove <artifact>` with NO proof about it is VACUOUS too** (exit 6,
326
+ `no proof is about <ref> — dt list proofs --missing`). It used to answer `proofs: 0 passed · …` at
327
+ exit 0 — and the orientation block tells every session to quote this command before saying an
328
+ artifact works, so a green result from a question nobody had asked was rule 7's own failure mode
329
+ shipped as a feature. `--all` over a workspace with no proofs keeps its exit 0: "run everything" is
330
+ truthfully green at zero; "prove THIS artifact" is not.
331
+
332
+ **UNAVAILABLE is checked before the fixture**, so a machine that cannot answer never reports
333
+ NO-FIXTURE — which would read as a fact about the workspace's data rather than about this laptop.
334
+ A `writes` proof with no fixture is UNAVAILABLE too: the artifact is fine, and this invocation
335
+ cannot answer for it without `--here`.
336
+
337
+ ⚠ **`--strict` means two different things, on purpose.** On `dt prove --all` (and the artifact form)
338
+ it makes **UNAVAILABLE fatal** — a board is what a hook runs, and "this machine could not ask" is a
339
+ gap the hook may want to fail on. On `dt status` it fails on a **FAIL tail** in the ledger and says
340
+ nothing about UNAVAILABLE, because that line reports what this machine has already proved rather
341
+ than running anything.
342
+
343
+ ## reading the surface
344
+
345
+ ```bash
346
+ dt prove <proof> # one proof: the transcript, and one of six codes
347
+ dt prove skills/<id> # every proof whose `about` names this artifact — a board;
348
+ # exit 6 when NOTHING is about it
349
+ dt prove --all [--kind gate|live] [--external] [--strict] [--json]
350
+ dt list proofs [--missing] [--filter …] [--json]
351
+ dt get proofs/<id> # the record, plus availability and the ledger tail
352
+ dt get commands/<id> # the artifact, plus the proofs that are about it
353
+ dt status [--strict]
354
+ ```
355
+
356
+ - **`--all` and the artifact form are BOARDS: ONE line per proof**, then a summary
357
+ (`proofs: 3 passed · 1 failed · 0 unavailable · …`). A transcript is what a single-proof run is
358
+ for. A proof with a `perform` step is **listed, never started**, so a board can never exit 5 —
359
+ "one of your forty proofs would like a human" is not an answer a hook can act on. A proof that
360
+ THROWS is that proof's FAIL, with a ledger row, not a silently green skip. The `writes` proof with
361
+ no fixture is not an exception to that: it SETTLES `UNAVAILABLE` with a row, in the single-proof
362
+ form and on the board alike (and `--strict` is what makes it fatal).
363
+ - **`--strict`** makes UNAVAILABLE fatal on a board. It is a flag rather than the default because a
364
+ proof needing a credential is ordinarily unavailable on a cloud session.
365
+ - **`dt list proofs`** appends two COMPUTED columns no record carries: `availability` on THIS machine
366
+ (with the fix, never a `.env` value) and `last`, the ledger tail (`PASS 2026-09-07 [notes/b]`, or
367
+ `never`). **`--missing`** inverts it — one line per artifact no proof is about — and takes no
368
+ filter, because it lists artifacts rather than proofs.
369
+ - **`dt compile` prints a coverage line on EVERY compile**, even at zero:
370
+ `proofs: 4 declared · commands 1/2 · skills 0/1 · scripts 1/1 · bindings 0/0`. And it **nudges
371
+ once** per NEW command or script with no proof, naming the file to write. `--missing` is the same
372
+ question answered by name.
373
+ - **`dt status`** counts each proof's LAST verdict on this machine
374
+ (`proofs: 7 declared · 1 passed · 1 failed · 1 unavailable · 4 never`) and `--strict` exits 1 when
375
+ any tail is a FAIL.
376
+
377
+ ## the eval layer — proving a skill actually TEACHES
378
+
379
+ A skill's proof can assert that its file compiles, that the script it names runs, that the record it
380
+ promises appears. None of that answers the question a skill exists for: does a fresh session FIND
381
+ it, LOAD it and DO the job right. That is answered by running sessions and scoring them. **The
382
+ engine never runs an agent** — this is a procedure, and the record of it belongs wherever the
383
+ workspace keeps its findings, not in a `proofs` record.
384
+
385
+ The method, as run on 2026-09-05 against a rebuilt orientation block:
386
+
387
+ 1. **Pick five REAL tasks** out of this workspace's own history — things that were actually done,
388
+ across different modules and different grains. Invented tasks measure how well you invent.
389
+ 2. **Give each session ONLY the artifact under test** plus whatever the harness injects anyway. No
390
+ briefing, no hints, no follow-up steering. Five separate blind sessions, one task each.
391
+ 3. **Have each session plan the task**, then **verify its own plan against the records and skills
392
+ that actually did the work**, then **review the packet adversarially** — what misled it, what it
393
+ could not find, what it invented.
394
+ 4. **Score each session on the sheet below**, and read the failures for their SHAPE rather than
395
+ one at a time. Five sessions failing at five different points is noise; five failing at the same
396
+ depth is the finding.
397
+ 5. **Repair, then rescore.** A repair nobody re-measured is a hypothesis.
398
+
399
+ | dimension | 1 | 3 | 5 |
400
+ |---|---|---|---|
401
+ | **found it** | never opened the artifact | opened it after guessing wrong first | loaded it before acting |
402
+ | **routed the noun** | wrote to the wrong collection | right collection, wrong grain | right collection, right grain |
403
+ | **respected the refusals** | proposed a write the store would reject | hit one refusal, recovered | wrote nothing that would be refused |
404
+ | **followed the order** | ignored a stated ordering constraint | followed it after being wrong once | followed it first time |
405
+ | **joined correctly** | missed a join the data requires | found it by reading records | found it from the artifact's own prose |
406
+ | **invented nothing** | invented a field, a status or a path | invented one, caught it | cited every value it used |
407
+
408
+ Two findings worth carrying, because they generalize: a `use_when` clause earns its tokens exactly
409
+ where it encodes an **order** or a **refusal**, and loses them where it paraphrases the description;
410
+ and every session failed at the same depth — a JOIN, an EXECUTABLE nothing rendered, or a REFUSABLE
411
+ WRITE (a closed enum, a required field) the prose never showed.
412
+
413
+ ## common mistakes
414
+
415
+ | mistake | reality |
416
+ |---|---|
417
+ | a `count: { _gte: 1 }` on a collection that already has records | it holds before the step runs — that is VACUOUS, and `_delta` is what you meant |
418
+ | `about:` naming a skill that does not exist | compile fails: a proof about nothing reports success forever |
419
+ | a nested key meant as a field comparison | it is a reference HOP; compile refuses it by name |
420
+ | `record: notes/b`, or any target but `{record}` | the `given` picks the record; compile refuses a second target |
421
+ | expecting a `{record}` in a `run:` step to reach the shell as a reference | it does — that one IS substituted; it is every OTHER brace that passes through untouched |
422
+ | `where:` left bare, or `where: {}` | it asserts nothing and the proof passes having measured nothing — refused at compile |
423
+ | judging a `writes` proof by running it with `--here` on real records | that is what a fixture and a sandbox are for; `--here` is the exception, not the default |
424
+ | expecting `--all` to run the `perform` proofs | it lists them; a board is for a hook, and a hook cannot perform |
425
+ | reading a step's stdout as a tail | it is captured and judged WHOLE, to 64 KB — over that, write a file and assert `path:` |
426
+ | `stdout: hello` — a scalar where a filter goes | refused at compile; a bare `stdout:`, a number and a list are the same defect — write `{ _contains: 'hello' }` |
427
+ | committing the ledger | it is per-machine evidence under gitignored build output; the SOURCE is what travels |
428
+ | a fixture carrying a `package.json` or a descriptor | only `data/` is admitted — otherwise a proof can bring the rules it is judged by |
429
+ | a proof whose descriptor is not committed | a sandbox is cut from HEAD; commit the schema before the proof that needs it |
430
+ | treating UNAVAILABLE as a failure | this machine cannot answer the question — the artifact is not implicated |
431
+ | a `kind: eval` proof | there is no such kind; the eval layer is a procedure, and the engine never runs an agent |
432
+ | one `expect` row carrying two forms' keys | refused at compile — a row judged as the wrong form asserts nothing and passes |
433
+ | `{record}` in `given.where` | nothing substitutes there; the given is what picks the record |
434
+ | a relative `path:` read as "beside where I typed dt" | it is resolved against the WORKSPACE root — the sandbox's, inside a `writes` proof |
435
+ | `dt prove skills/x` exiting 0 as proof that the skill works | exit 6 when nothing is `about` it; check the code, not the absence of red |
@@ -160,6 +160,14 @@ Skills ship with modules and are read by any operator on any machine:
160
160
  renamed verb or a changed limit in a skill sends every future session down the old path
161
161
  confidently. When you catch a skill lying, fixing it is part of the task you are on, not a
162
162
  follow-up.
163
+ - **Give it a proof, and know what a proof cannot cover.** A `proofs` record (`proofs.md`) pins the
164
+ mechanical half: the script the skill names runs, the record it promises appears, the path it
165
+ files to exists. Write it in the same commit — `dt add skills` nudges you with the path the moment
166
+ it writes the skill (compile's own nudge covers new commands and scripts, never skills), and
167
+ `dt list proofs --missing` names every artifact nobody claimed anything about. ⚠ **A
168
+ green proof is not evidence the skill TEACHES.** Whether a fresh session finds it, loads it and
169
+ does the job right is the eval layer: real tasks, blind sessions, a scoring sheet — a procedure,
170
+ never something the engine runs.
163
171
  - **Retire what nothing loads.** A skill nobody uses still costs its line in every session's index.
164
172
  Fold it into a sibling or delete it; git keeps the text.
165
173