@hviana/sema 0.7.2 → 0.7.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (43) hide show
  1. package/AGENTS.md +95 -843
  2. package/README.md +11 -11
  3. package/dist/src/mind/mind.js +14 -0
  4. package/dist/src/store-sqlite.js +17 -0
  5. package/dist/src/store.d.ts +51 -4
  6. package/dist/src/store.js +81 -14
  7. package/docs/INDEX.md +71 -0
  8. package/docs/INVARIANTS.md +19 -0
  9. package/docs/architecture/bounded-reads.md +85 -0
  10. package/docs/architecture/caches.md +89 -0
  11. package/docs/architecture/commonality.md +45 -0
  12. package/docs/architecture/cost-model.md +71 -0
  13. package/docs/architecture/determinism.md +73 -0
  14. package/docs/architecture/exact-vs-approximate.md +47 -0
  15. package/docs/architecture/factored-machinery.md +28 -0
  16. package/docs/architecture/fold-contract.md +87 -0
  17. package/docs/architecture/halo-sketch.md +99 -0
  18. package/docs/architecture/match-project.md +62 -0
  19. package/docs/architecture/mechanism-market.md +95 -0
  20. package/docs/architecture/memoization.md +96 -0
  21. package/docs/architecture/meter.md +55 -0
  22. package/docs/architecture/saturation.md +92 -0
  23. package/docs/architecture/store.md +79 -0
  24. package/docs/architecture/thresholds.md +79 -0
  25. package/docs/failures/tempting-but-wrong.md +144 -0
  26. package/docs/harness/gates.md +56 -0
  27. package/docs/mechanisms/alu.md +75 -0
  28. package/docs/mechanisms/cast.md +75 -0
  29. package/docs/mechanisms/confluence.md +36 -0
  30. package/docs/mechanisms/cover.md +54 -0
  31. package/docs/mechanisms/extraction.md +53 -0
  32. package/docs/mechanisms/prefix-completion.md +54 -0
  33. package/docs/mechanisms/recall.md +69 -0
  34. package/docs/mechanisms/reference.md +58 -0
  35. package/jsr.json +1 -1
  36. package/package.json +1 -1
  37. package/src/mind/mind.ts +14 -0
  38. package/src/store-sqlite.ts +19 -0
  39. package/src/store.ts +92 -16
  40. package/test/89-completion-recursion.test.mjs +30 -10
  41. package/test/96-bytes-walk-termination.test.mjs +115 -0
  42. package/test/97-store-seed.test.mjs +105 -0
  43. package/HOW_IT_WORKS.md +0 -5836
package/AGENTS.md CHANGED
@@ -1,895 +1,147 @@
1
1
  # AGENTS.md — the Sema development manual
2
2
 
3
- The working manual for anyone (human or AI agent) changing Sema. It is organized
4
- around the **engineering patterns the codebase runs on**: what each pattern is,
5
- where it is applied, how to follow it, and what breaks if you don't. You should
6
- be able to develop against this document alone; read
7
- [HOW_IT_WORKS.md](HOW_IT_WORKS.md) only when you need the theory behind a
8
- pattern, not to get work done.
9
-
10
- ---
3
+ The working manual for anyone (human or AI agent) changing Sema. For pattern
4
+ detail, see `docs/INDEX.md` → `docs/architecture/*.md`. You should be able to
5
+ develop against this document and docs/ alone; read the theory in
6
+ `docs/architecture/` when you need why a pattern holds, not to get work done.
11
7
 
12
8
  ## 1. Orientation
13
9
 
14
10
  Sema is a deterministic reasoning engine with no ML runtime: a content-addressed
15
- graph store (the knowledge), two approximate vector indexes over it (the search
11
+ DAG store (the knowledge), two RaBitQ-IVF vector indexes over it (the search
16
12
  accelerators), and a cost-based search that composes answers from stored facts
17
13
  (the inference). Everything is plain TypeScript, CPU-only, with `node:sqlite` as
18
- the only runtime dependency of the library.
14
+ the only runtime dependency.
19
15
 
20
16
  ```bash
21
- npm install # dev tooling + the parquet reader used by one example
17
+ npm install # dev tooling + parquet reader used by one example
22
18
  npm run build # tsc → dist/
23
19
  npm test # tsc && node --test test/**/*.test.mjs
24
20
  npm run demo # example/demo.ts — the four-note README demo
25
21
  ```
26
22
 
27
- Hard facts you must not fight:
23
+ Hard facts:
28
24
 
29
- - **Node ≥ 22.5** (`node:sqlite`). No native add-ons. If you want a GPU, you are
30
- working against the architecture.
31
- - **ES modules with explicit `.js` extensions** in imports. Keep the convention
32
- in new files.
25
+ - **Node >=22.5** (`node:sqlite`). No native add-ons.
26
+ - **ES modules with explicit `.js` extensions** in imports. Keep the convention.
33
27
  - **Determinism is the product.** Same seed + same deposit order + same query ⇒
34
- byte-identical answer. Tests pin it.
35
- - **Perception is a pure function of the bytes**, and the deposit and inference
36
- paths must compute the same tree for the same input (2.15). Anything that
37
- makes one side impose structure the other does not is a correctness bug, not a
38
- tuning choice.
28
+ byte-identical answer.
39
29
 
40
- The mental model, top to bottom:
30
+ Mental model, top to bottom:
41
31
 
42
32
  ```
43
- mind/pipeline.ts the grounding decider: mechanisms compete on one cost scale
44
- mind/mechanisms/* cover · cast · confluence · extraction · reference ·
45
- recall · prefix-completion · alu
46
- mind/* shared machinery: match/project, attention, recognition,
47
- junction ascent, graph search, learning, rationale;
48
- the substitution bridge lives beside them, not inside
49
- recall.ts
50
- store.ts AbstractStore: ALL domain logic of the DAG store
33
+ mind/pipeline.ts grounding decider: mechanisms compete on one cost scale
34
+ mind/mechanisms/* cover · cast · confluence · extraction · reference · recall · prefix-completion · alu
35
+ mind/* match/project, attention, recognition, junction ascent, graph search, learning, rationale
36
+ store.ts AbstractStore: ALL DAG store domain logic
51
37
  store-sqlite.ts the one concrete backend (thin SQL wrappers)
52
- geometry.ts + vec/alphabet/sema/canon
53
- vectors, the fold (content-defined cuts + two-ended
54
- seats), every derived threshold, and the injected
55
- content canonicalizer
56
- derive/ · alu/ · rabitq-ivf/
57
- firewalled sublibraries with their own READMEs and tests
38
+ geometry.ts + vec/alphabet/sema/canon vectors, fold, every derived threshold, canonicalizer
39
+ derive/ · alu/ · rabitq-ivf/ firewalled sublibraries (own READMEs, own tests)
58
40
  ```
59
41
 
60
- Lower layers never import higher ones. The three sublibraries import nothing
61
- from the rest of Sema (the ALU may use `../bytes.ts` only); Sema reaches them
62
- through narrow interfaces (`DeductionSystem`, `PipelineMechanism`,
63
- `VectorDatabase`). A change that makes a sublibrary import mind code is
64
- architecturally wrong, full stop.
65
-
66
- ---
67
-
68
- ## 2. The patterns
69
-
70
- Each subsection: what the pattern is → where it is applied → how to follow it.
71
-
72
- ### 2.1 Determinism as a contract
73
-
74
- No `Math.random`, no `Date.now` in behaviour, no iteration over unordered
75
- collections where order can reach output. All randomness flows from the config
76
- `seed` (alphabet, keyring, index PRNGs). Every tie-break bottoms out in a fixed,
77
- corpus-determined ordering — insertion order or lowest node id, with
78
- **first-inserted** as the universal no-evidence fallback (last-inserted was once
79
- used in one place; it was a bug).
80
-
81
- _Follow it:_ when you introduce any choice among equals, pick the tie-break
82
- explicitly and make it corpus-determined. If a test becomes flaky, you broke
83
- this contract, not the test.
84
-
85
- ### 2.2 Derived thresholds — never tuned
86
-
87
- Every similarity/decision threshold is a **formula** over the vector dimension
88
- D, the perception window W, or the corpus size N, defined once in
89
- `src/geometry.ts` (`mergeThreshold`, `identityBar`, `reachThreshold`,
90
- `significanceBar`, `estimatorNoise`, `conceptThreshold`, `consensusFloor`,
91
- `dominates`, …), with the corpus-scale readings beside them in
92
- `mind/traverse.ts` (`corpusN`, `hubBound`, `hubCap`, `atomReach`, `atomIsHub`).
93
- `src/config.ts` holds **capacities and budgets only** (cache byte budgets, batch
94
- sizes, vector-index parameters, query k, ALU precision, the seed).
95
-
96
- _Follow it:_ if you are about to add a tunable cutoff to config, derive it in
97
- `geometry.ts` instead. A threshold knob is a design bug here — one was already
98
- removed once. When a decision needs a scale, express it in D, W, or N.
99
-
100
- Two derivations bite in practice and are worth knowing before you reach for a
101
- bar:
102
-
103
- - `identityBar(D, W, len)` is the SCALE-AWARE identity claim (`1 − W/len`,
104
- floored at `mergeThreshold`). Reuse `mergeThreshold` for a whole-span claim
105
- and long spans silently tolerate whole windows of foreign content.
106
- - A bar calibrated for one quantity does not transfer to another. `chooseNext`
107
- carries a comment recording exactly this: the consensus floor is priced for
108
- pooled, N-scaled climb votes, and gating an N-invariant support count against
109
- it fails once N is large enough (test/40 pins it).
110
-
111
- ### 2.3 Exact decides, approximate proposes
112
-
113
- Every score from the vector indexes (`resonate`, `resonateHalo`) is a RaBitQ
114
- _estimate_. Identity is decided only by content-addressed lookup (`resolve`,
115
- `findLeaf`, `findBranch`, and the canonical fallback `canonResolve`), never by
116
- `score >= threshold`. Scores rank candidates and gate broad regions; bytes make
117
- decisions. Even the echo decision in `recall.ts` re-folds the top hit's bytes
118
- rather than trusting the estimate it already has.
119
-
120
- The same principle appears as **graded evidence ladders** — exact tier first,
121
- distributional second, geometric last — in five places built on one shape:
122
-
123
- - `resolve` in `mind/primitives.ts`: exact content-addressed fold →
124
- `canonResolve` (equivalence class, hash-then-verify).
125
- - `locate` in `mind/match.ts`: exact bytes → halo role → gist.
126
- - `alignGraded` in `mind/match.ts`: literal W-gram runs → halo-matched sites
127
- (the weave adds a third pass from the climb's own proposals — see
128
- `pipeline-mechanism.ts`).
129
- - `bridge` in `mind/resonance.ts`: junction containers by identity → edge
130
- junctions → synonym junctions → whole-gist resonance as last resort.
131
- - `crossRegionVotes` in `mind/attention.ts`: exact containers → single synonym →
132
- double synonym → `structuralResonance` (a synthetic gist; the one tier with no
133
- byte containment behind it, and gated hardest because of it).
134
-
135
- _Follow it:_ never reorder a ladder's tiers, and never let an approximate tier
136
- override an exact one. Two asymmetries in `attention.ts` encode that rule and
137
- must not be flattened: only the EXACT tier may explain ordinary votes away, and
138
- only container-backed evidence may consume its endpoints. If you need a new
139
- matcher, add a tier to the shared family (2.5), not a private score check.
140
-
141
- ### 2.4 One cost currency
142
-
143
- The graph search's cost ladder (`mind/graph-search.ts`, exported constants) is
144
- the single pricing scheme of the whole mind:
145
-
146
- ```
147
- MICRO (1e-3) advance over recognised material; per-byte unit of the A* heuristic;
148
- a RECOMPOSED form's onward edge
149
- STEP (1) follow one learned edge (EVERY hop, first or fifth); one computed
150
- result; one projection
151
- CONCEPT (10) a halo-mediated act (synonym hop, consensus climb); also the price
152
- of ABANDONING an edge chain early (graph-search's stop-here rule)
153
- PASS (1000/byte) carry a byte nothing explains
154
- ```
155
-
156
- Only the _ordering_ matters. The pipeline weighs whole mechanisms in the same
157
- units: `weight = moves + PASS · unaccounted-bytes`, so a mechanism-level choice
158
- and a byte-level choice are the same kind of decision. Weights are compared at
159
- STEP resolution (`grade = ⌊w/STEP⌋`); at equal grade the candidate reporting
160
- fewer `scaffolding` bytes wins, and only then does the mechanism list's order
161
- decide.
162
-
163
- Two pricings inside `graph-search.ts` are easy to "simplify" and are not free to
164
- change: charging every edge hop STEP is what makes the lightest derivation the
165
- SHORTEST chain (charging later hops nothing made every stopping depth tie), and
166
- the stop-here rule at CONCEPT-above-chain-cost is what keeps a genuine fixpoint
167
- preferable to giving up at the same depth.
168
-
169
- _Follow it:_ place any new cost deliberately in the ordering; never make the A\*
170
- heuristic exceed a real per-byte cost (admissibility breaks silently — answers
171
- degrade, nothing errors). Never encode _policy_ as cost: "computation always
172
- wins" is implemented by masking colliding sites in cover, not by pricing. Keep
173
- policy in callers, the engine neutral.
174
-
175
- ### 2.5 One factored machinery: match → project, under a gate
176
-
177
- `mind/match.ts` is the shared family every generalising mechanism configures:
178
- matchers (`locate`, `alignRuns`, `alignGraded`, `alignAround`/`frameSlots`,
179
- `bestHaloMate`, `haloSiblings`, `analogyStrength` and its structural tier
180
- `sharedFrameStrength`, `spanHalo`, `spanSynonymStrength`), projections
181
- (`follow`, `reverseContext`, `project`, `conceptHop`), the span-shape family
182
- (`skillExemplar`, `isSpanShaped`, `containsSpan`), and the STRUCTURAL gates that
183
- are byte predicates rather than derived thresholds (`isSpanShaped`,
184
- `carriesFillers`, `voicesDisplacedFiller`). `mind/traverse.ts` owns the graph
185
- readings (`edgeAncestors`, `reachOf`, `chooseNext`/`chooseAmong`, `guidedFirst`,
186
- `leadsSomewhere`, `allWindowsAreScaffolding`, `formsOpenedBy`) and the corpus
187
- scale (`corpusN`, `hubBound`, `hubCap`, `atomReach`).
188
-
189
- _Follow it:_ before writing a new generalising mechanism, express it as a
190
- (matcher, direction, gate) triple. If those already exist, the mechanism is a
191
- configuration — write only the configuration. If it genuinely needs a new
192
- matcher or projection, add it **to the shared family** with a derived gate
193
- (2.2), never as a private helper. A mechanism file that re-implements locating,
194
- aligning, edge-following to a fixpoint, predecessor-picking, or fan-out capping
195
- is reintroducing duplication that was deliberately removed.
196
-
197
- A shared analysis must not live inside a mechanism. If `pipeline-mechanism.ts`
198
- (the shared contract and `Precomputed`) or a post-grounding stage has to import
199
- _out of_ `mechanisms/`, the dependency is inverted and the market's decoupling
200
- (2.6) is broken — deleting that mechanism would break the shared container. The
201
- span-shape family (`isSpanShaped` / `containsSpan` / `skillExemplar`) lives in
202
- `match.ts` for exactly this reason: its two consumers reach it without knowing
203
- extraction exists. `alignAround` is there on the same grounds — the substitution
204
- bridge and the frame reading both need the same gaps, and ask opposite questions
205
- of them (the bridge EXPANDS a gap until the query side attests, because a
206
- substitution claims equivalence; the frame reading CONTRACTS it to its varying
207
- core, because a reference claims only position). Neither reading derives the
208
- other, and one aligner serves both. `formsOpenedBy` (`traverse.ts`) is the
209
- retrieval counterpart: "which trained forms does this byte run open?" is a
210
- question about the STORE, so it sits with the graph readings rather than inside
211
- the mechanism that first needed it.
212
-
213
- The **frame reading** (`alignAround` / `contractGap` / `frameSlots` /
214
- `carriesFillers`, plus `Precomputed.frames`) is the worked example of this
215
- pattern at full length. Sema is otherwise fully GROUND — nothing anywhere
216
- represents a position whose occupant comes from the context rather than the
217
- corpus — so without it no mechanism can tell "the corpus does not explain these
218
- bytes" (PASS, refuse) from "these bytes occupy a place the corpus keeps open"
219
- (bind). Split along the §2.5 triple the notion lands in three places, each at
220
- its own altitude:
221
-
222
- - the **matcher** (`frameSlots`) REPORTS and does not judge: every place a
223
- pairing varies, contracted to its varying core, tagged
224
- substitution/insertion/deletion, plus how much the two share. It rejects
225
- nothing, so a consumer can apply its own reading to a pairing another consumer
226
- would throw away.
227
- - the **gate** (`carriesFillers`) is the much stronger claim that a slot may be
228
- VOICED through, so it is deliberately not folded into the matcher: a consumer
229
- taking the matcher's answer as permission to voice would be making exactly the
230
- claim the licence withholds.
231
- - the **inventory** (`Precomputed.frames`) elects no frame, because a slot is a
232
- property of a PAIRING, not of the query. Election is each consumer's own —
233
- `reference.ts` keeps the modal slot signature, and a consumer wanting another
234
- reading is not fighting that one.
235
-
236
- **Voicing gates belong to the consumer that voices, never to the matcher.** The
237
- four reference applies — the frame must dominate the query, each slot must reach
238
- one window on both sides, an insertion or deletion disqualifies the pairing,
239
- fillers must be pairwise distinct — are all requirements for substituting and
240
- SPEAKING, not for knowing where a pairing varies. Inside `frameSlots` they make
241
- the shared reading useless to anyone else: measured over four real pairings,
242
- only reference's own survives, while a definite description standing where a
243
- proper noun stands, a pure insertion and a sub-window difference all come back
244
- as NOTHING. Nothing fails to compile and no test notices.
245
-
246
- _Follow it:_ a shared analysis with exactly ONE consumer is unproven, whatever
247
- its address. Before declaring machinery shared, run a second consumer's real
248
- case through it and check the answer is not `null`. If every gate you wrote
249
- happens to be one your own mechanism needs, they are not the matcher's gates.
250
-
251
- Making a notion available is not the same as imposing it, and two mechanisms
252
- deliberately do **not** consume this one: the substitution bridge (its
253
- substitution asserts equivalence, and it grounds through its candidate's
254
- continuation UNSUBSTITUTED, so admitting a slot-gap there voices the corpus's
255
- filler for the asker's referent) and CAST (its frame gate is weave-local while a
256
- slot is cohort-local — §2.7 again).
257
-
258
- Related single-definition contracts (define once, import everywhere):
259
-
260
- - `contentLevels` (`geometry.ts`) — the ONE boundary rule: where a stream
261
- segments and at what level. `contentBoundaries` is a projection of it, not a
262
- second copy; it used to carry its own rolling-hash loop, which is exactly how
263
- a write side and a read side drift apart without a type error.
264
- - `canonical.ts` — the write/read contract for canonical segmentation
265
- (`canonicalWindows`, `chainReach`, `leafIdRun`, `leafIdPrefix`, `windowIds`).
266
- Learning writes through it; recognition, attention, confluence, the bridge and
267
- prefix completion read through it. Changing one side means changing this file
268
- — drift between sides breaks canonical recognition with **no type error**.
269
- - `junction.ts` — the content-addressed "which learnt whole contains these two
270
- forms?" ascent, shared by the bridge and cross-region attention, with its
271
- per-response `WalkCache` and once-per-candidate seed computation.
272
- - `joinWithBridge` (`resonance.ts`) — the one out-of-search way to join two
273
- answer spans; it emits a `bridgeMiss` trace step on a bare join.
274
- - `dismissedKnownContent` (`bridge.ts`) — the one IGNORED-KNOWN test ("does the
275
- unaccounted remainder contain a STORED window?"), shared by the substitution
276
- bridge's own acceptance and CAST's frame-tier comparison gate.
277
- - `sharedReachMemo` (`traverse.ts`) — the one definition of the ancestor-reach
278
- memo's lifetime (session-scoped between writes, cold under a trace). There
279
- used to be two memos that never met.
280
- - `guidedFirst` (`traverse.ts`) — the one answer-shaped "what does this lead
281
- to?" read (guided pick merged with the first-inserted fallback).
282
- - `leadsSomewhere` (`traverse.ts`) — the one admission predicate for recognition
283
- sites (edge-or-halo, via existence probes).
284
- - `isChunk` (`sema.ts`) — the one "children are all leaves" predicate.
285
- - `twoEndedSeat` (`sema.ts`) — the one positional-coordinate algebra, shared by
286
- perception, `fold`, and every synthetic/canonical fold.
287
-
288
- ### 2.6 The mechanism market (the free-will architecture)
289
-
290
- Every grounding mechanism — including the ALU and user extensions — implements
291
- the same interface, `PipelineMechanism` (`mind/pipeline-mechanism.ts`): optional
292
- `parse` (authoritative computed spans, collected before anything else), `floor`
293
- (an admissible lower bound, or `null` when the mechanism structurally cannot
294
- fire), and `run` (candidate answers). The decider in `mind/pipeline.ts`
295
- (`think`) holds a plain list (`defaultMechanisms`: cover, cast, confluence,
296
- extraction, reference, recall, prefix-completion, plus the ALU and any user
297
- mechanisms) and never branches on which mechanism it is holding.
298
-
299
- Four constraints make the market honest — verify all four for anything you add:
300
-
301
- 1. **Decoupled.** Zero cross-imports between mechanism files. Adding one never
302
- touches another. No mechanism asks "did an extension already decide?" — it
303
- only asks "can I still beat the incumbent?".
304
- 2. **Declared competence.** Gates are binary structural preconditions checked
305
- inside `floor`/`run` (query length, anchor shape, weave existence), never
306
- learned scores — so the rationale states exactly why a mechanism abstained.
307
- 3. **Visible budget.** Every corpus-scale loop is capped at a named constant
308
- (`√N` via `hubBound`, `k = 2·recallQueryK`), enforced at the store level
309
- (2.8).
310
- 4. **Evidence travels.** Every candidate carries `accounted` (query spans its
311
- structural evidence explains), `moves` (its acts, priced on the ladder), and
312
- `unexplained` (a diagnostic label). Two optional fields let a mechanism state
313
- things only it can know: `scaffolding` (answer bytes lifted from spans
314
- nothing recognised — the equal-grade tie-break) and `complete` (this answer
315
- is a trained form's own continuation reached through an identity claim about
316
- the query, so post-grounding must not extend it). The decider honours both
317
- without ever asking which mechanism set them. It sees only weights.
318
-
319
- Two disciplines inside the loop:
320
-
321
- - **Admissible-floor pruning.** `floor` runs for every mechanism in list order,
322
- before any `run`; `run` fires only if the floor can still beat the incumbent
323
- (`worthRunning`). Cover runs first so a computed span's near-zero cost prunes
324
- everything after it through this same mechanism — not a special case.
325
- - **Investment discipline.** `worthRunning` is also passed _into_ `floor`: a
326
- floor that would first-touch an expensive shared analysis (the climb, the
327
- weave) checks its cheapest possible bound against the incumbent _before_
328
- paying, and returns the uninvested bound when it already loses. Never compute
329
- a shared analysis just to discard it.
330
-
331
- Evidence accounting rules that bite:
332
-
333
- - **Read-out content is selectively accounted.** Extraction's located frames are
334
- always evidence; the span between them counts only when _both_ borders were
335
- located. An open-ended read is content-novel and is priced by exclusion
336
- (PASS/byte), like the cover's carried literals.
337
- - **Reverse reading is not derivation.** A `reverseContext` projection produces
338
- bytes but explains nothing forward: `accounted = []`, weight ≈ PASS·|query|.
339
- It is the designated last resort by arithmetic, not by rule.
340
- - **An act you PAID for is accounted.** The mirror of the rule above: the
341
- bridge's corroborated substitutions cost a CONCEPT each in `moves`, so leaving
342
- their spans unaccounted charges the same act twice — and the PASS-per-byte
343
- charge is far the larger (measured: a bridge matching 28 of 29 bytes declared
344
- the whole query unexplained and lost).
345
- - `accounted` is a COST-LADDER quantity, not a coverage one. `cover.ts`
346
- deliberately leaves masked computed spans out of it so PASS-bridged bytes are
347
- still charged, so a fully-explained query can report `accounted: []`. The
348
- post-grounding fusion gate therefore reads `accounted ∪ pre.computed`, not
349
- `accounted` alone.
350
- - `unexplained`, `narrowDecision`, and `thinGrounding` are **observational
351
- only** — they appear in the trace and never alter the decision.
352
-
353
- ### 2.7 Two measures of commonality — pick the right population
354
-
355
- "Is this content discriminative?" has two formally independent answers, and
356
- using the wrong one is a semantic bug the type system cannot catch:
357
-
358
- - **Corpus-global** — reference set: all learned contexts. Tooling: `reachOf` +
359
- `dominates` (+ `corpusN`). Used by the climb's IDF weighting and confluence's
360
- filler/scaffolding gate. Answers "does this discriminate anything in the
361
- store?"
362
- - **Weave-local** — reference set: the structures aligned with _this query_.
363
- Tooling: the `depth[]` array built in `computeWeave` + `MIN_WEAVE` +
364
- `dominates`. Used by CAST's frame gate. Answers "does this discriminate among
365
- the structures this query activates?"
366
-
367
- `depth[]` counts **distinct covering structures**, not accumulated alignment
368
- weight: the frame test compares it against a COUNT of aligned points, so
369
- accumulating weight there compares weight-mass against a cardinality. It reads
370
- like a harmless refinement and inverts the frame verdict (measured: 29 of 42
371
- bytes reading FRAME against 6 of 42, with nothing else changed).
372
-
373
- _Follow it:_ when adding a gate on "shared vs. discriminative", write down which
374
- population your question is about before choosing the tool. Substituting one for
375
- the other in CAST misfires on reordered single-fact queries (test 17 pins this).
376
-
377
- ### 2.8 Bounded reads — the cap lives in the store
378
-
379
- No per-query read may grow with the corpus. The cap is √N (`hubBound(ctx)`,
380
- derived from `corpusN(ctx)` — the one definition of corpus size, floored at 2).
381
- Crucially, the cap is enforced **at the store level**:
382
-
383
- - LIMITed reads: `nextFirst`, `prevFirst`, `parentsFirst`, `containersSlice`. In
384
- an adapter these must be real `LIMIT ?` statements — never "materialise then
385
- slice". Reading `hubBound + 1` parents decides "hub or not" exactly.
386
- - Existence probes: `hasNext`, `hasParents`, `hasContainers`, `hasHalo`,
387
- `prevCount` — indexed point probes that never decode vectors or unpack blobs.
388
- Use them for every "does this lead anywhere?" question instead of
389
- `next(id).length > 0`.
390
- - Prefix-capped reads: `bytesPrefix(id, cap)` and `contentLen(id, cap)`. A
391
- candidate that exceeds the cap is rejected without reconstructing it — the
392
- weave, the junction walks and the bridge all read this way, and uncapped reads
393
- there cost seconds per query on a large store.
394
- - `chainRun` climbs transparent scaffolding chains in one bounded read.
395
- - The full materialising reads (`next`, `prev`, `parents`, `containers`) exist
396
- for maintenance and inspection only. Keep them off hot paths.
397
-
398
- `edgeAncestors` is the reference consumer: it decides saturation from LIMITed
399
- reads alone, by five named stops (predecessor fan-in, distinct-context limit,
400
- parent fan-out, the cumulative **lateral-cone** bound, and **byte-atom**
401
- commonality). The last two are the ones a new walk forgets. An atom carries no
402
- kid/contain rows by construction, so its commonality is unmeasurable and must
403
- not default to "maximally rare" — `atomReach`/`atomIsHub` are the honest floor.
404
-
405
- _Follow it:_ any new fan-out walk uses `hubBound`/`hubCap` — do not invent a
406
- second convention, and do not call `edgeSourceCount()` or
407
- `Math.ceil(Math.sqrt(...))` inline.
408
-
409
- ### 2.9 Template-method store
410
-
411
- `store.ts` (`AbstractStore`) owns **all** domain logic: exact dedup,
412
- byte-verified near-dedup, lazy gist indexing and bridge promotion, halo
413
- quantization and exact in-session accumulators, containment buffering, write
414
- batching, LRU budgets, compaction cadence. `store-sqlite.ts` implements only the
415
- abstract `_db*`/`_vec*` methods as thin statement wrappers.
416
-
417
- _Follow it:_ a new backend subclasses `AbstractStore` and implements the
418
- abstract methods — nothing else. If you find yourself re-implementing dedup or
419
- indexing logic in an adapter, stop. Facts an adapter (and any store caller) must
420
- respect:
421
-
422
- - Branch ids are dense non-negative integers minted in order, never deleted.
423
- Single-byte leaves are **implicit negative ids** (−256…−1) with no DB row —
424
- id-iterating code must handle both ranges.
425
- - Flat branches (all-leaf children) are stored as raw bytes with an empty kids
426
- blob as marker (`flatKidsBytes`/`flatBytesKids`).
427
- - `bytes()`/`bytesPrefix()` return arrays **shared with caches — never mutate**;
428
- copy first.
429
- - `contentLen(id, cap?)` reads a node's byte length; pass `cap` when exact
430
- length beyond a bound doesn't matter.
431
- - The **canon index** (`canonAdd`/`canonFind`/`canonCount`/`eachContent`) is an
432
- OPTIONAL capability: a backend may omit all four, and resolution then simply
433
- has no equivalence fallback. The store never learns what the equivalence IS —
434
- the canonicalizer is injected by the caller and every candidate is verified by
435
- re-canonicalizing its bytes, so a hash collision costs a read, never a wrong
436
- id.
437
- - Maintenance entry points (`compactContentIndex`, `repairContentIndex`,
438
- `Mind.buildCanonIndex`) are batch operations for checkpoints, never the hot
439
- path. `buildCanonIndex` is incremental — it remembers the last indexed id in
440
- store meta — and must be run under the SAME canonicalizer queries will carry.
42
+ Lower layers never import higher ones. Sublibraries import nothing from Sema
43
+ except `../bytes.ts` (ALU); Sema reaches them via `DeductionSystem`,
44
+ `PipelineMechanism`, `VectorDatabase`.
441
45
 
442
- ### 2.10 The async/sync seam and pre-resolution
46
+ ## 2. Invariants — routing table
443
47
 
444
- Perception, recognition, and the graph search are **synchronous**; anything
445
- touching the ANN indexes is **async**. A synchronous consumer that needs
446
- resonance uses _pre-resolution_: gather the async answers first (concept
447
- siblings, connectors, ALU operand meanings), hand them in as maps
448
- (`resolveConcepts`/`resolveConnectors` in `mechanisms/cover.ts` are the models).
449
- Do not try to make the search async.
48
+ Five invariants. Violate one and the system degrades silently — tests pin them.
450
49
 
451
- ### 2.11 Per-response memoization
50
+ | # | Invariant | Rule | Where it lives |
51
+ | - | ----------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------- |
52
+ | 1 | Determinism | No `Math.random`/`Date.now` in behaviour; all randomness from `seed`; tie-breaks are corpus-determined (insertion order / lowest id) | `docs/architecture/determinism.md` → `src/config.ts`, `src/alphabet.ts` |
53
+ | 2 | Derived thresholds | Every cutoff is a formula over D/W/N in `geometry.ts`; `config.ts` holds capacities and budgets only — no tunable thresholds | `docs/architecture/thresholds.md` → `src/geometry.ts` |
54
+ | 3 | Exact decides, approximate proposes | Vector scores (`resonate`) rank only; identity via content-addressed lookup (`resolve`/`findLeaf`/`canonResolve`); graded ladders are exact tier first | `docs/architecture/exact-vs-approximate.md` → `src/mind/primitives.ts`, `src/mind/match.ts` |
55
+ | 4 | One cost currency | Single ladder `MICRO`/`STEP`/`CONCEPT`/`PASS`; `weight = moves + PASS·unaccounted`; compare at `STEP` grade | `docs/architecture/cost-model.md` → `src/mind/graph-search.ts`, `src/derive/` |
56
+ | 5 | Bounded reads | No per-query read grows with N; cap is `hubBound = √N` enforced at the store via `LIMIT` reads, existence probes, and `bytesPrefix` caps | `docs/architecture/bounded-reads.md` → `src/store.ts`, `src/mind/traverse.ts` |
452
57
 
453
- Asking never writes, which is the only reason per-response memos are sound.
454
- `Precomputed` (`pipeline-mechanism.ts`) is the shared response-scoped container:
455
- eager fields (recognition, computed spans, guide, the evidence-breadth constant
456
- `k`) plus **lazily-cached methods** for expensive analyses (`attention()` — the
457
- consensus climb, `weave()`, `resonance()` — the response's ONE top-k
458
- content-index read, `frames()` — the frame/slot inventory,
459
- `spanShapedOf`/`spanShapedAll`, `queryWindows`, `queryResolved`, `windowsOf`,
460
- `reachMemo`) — each computed at most once, shared by mechanisms and
461
- post-grounding stages, and never computed if nobody asks. The async ones are
462
- cached **by promise**, so a second caller awaits the first computation rather
463
- than starting another.
464
-
465
- Mind-level memos (`climbMemo`, `recogniseMemo`, `perceiveMemo`, `canonMemo`,
466
- `_resolvedSubtrees`) are assigned in `beginResponse()` and nulled in
467
- `endResponse()` — a new per-response memo must be added to both. A conversation
468
- supplies its own maps for the first four, so they persist across turns.
469
- `_edgeChoice` is a Mind field CLEARED in `endResponse()` (not re-created).
470
- `_gistCache` is a SESSION-lifetime Mind field (32 MB, `mind.ts`) never touched
471
- by begin/endResponse: a node's bytes are immutable and perception is pure, so a
472
- cached gist is valid for the store's lifetime.
473
-
474
- **Which memos a trace bypasses, and why the answer is "almost none".** Only
475
- `_edgeChoice` (via `guidedNext`) and `sharedReachMemo` are trace-bypassed —
476
- there, a memo hit would swallow a repeat's `disambiguate` step or black out
477
- reach detail the trace serialises. `perceiveMemo`, `recogniseMemo` and
478
- `climbMemo` are **always** consulted, tracing or not, and that is a correctness
479
- contract rather than a speed choice: `recogniseImpl` walks the query tree
480
- through `foldTree`, whose subtree-resolution fast path skips `visit` — and
481
- therefore skips EMITTING SITES — for any subtree already cached. A second call
482
- on identical bytes is not idempotent; it finds strictly fewer sites (observed:
483
- 31 → 5). Bypassing these under trace made every traced turn re-run recognition
484
- at each of the many call sites that recognise the same query, each call more
485
- incomplete than the last, measurably changing which mechanism grounded the
486
- answer (test/42 pins this). Trace steps still fire on a cache hit, so a hit is
487
- never silent.
488
-
489
- _Follow it:_ an expensive analysis a new mechanism needs goes on `Precomputed`
490
- as a lazy method, not inside the mechanism. Do not add a `if (!ctx.trace)` guard
491
- to a memo without checking whether the memoised function is idempotent. And
492
- still **never benchmark with a trace attached** — the two bypassed memos, plus
493
- the trace's own allocation, measure a different machine.
494
-
495
- ### 2.12 Caches are budgets, not correctness
496
-
497
- Every in-memory acceleration structure is a `BoundedMap` with a byte budget; a
498
- miss re-derives from durable state. The degradation order is fixed: what is lost
499
- under pressure is always speed or reach (a re-perception, a duplicate probe,
500
- reduced resonance until repair), never identity, reconstruction, or a learned
501
- relation.
502
-
503
- The deposit-path caches (`_depositTrees`/`_depositLens` for folded segments,
504
- `_internIds` for already-interned tree nodes, `_resolvedSubtrees` for resolved
505
- ones) obey the same rule, with one extra obligation: `contentFoldIncremental`
506
- reuses segments keyed on their **offsets**, which cannot witness that the
507
- underlying bytes agree. The caller must discharge that structurally — the
508
- deposit cache is keyed by the prefix's own bytes, and a conversation's fold
509
- advances only by append. A caller that cannot make the same argument passes no
510
- `prev` at all; the cold path is always correct. Handing it a mismatched `prev`
511
- produced a wrong tree on 336 of 400 random streams.
512
-
513
- _Follow it:_ new caches get budgets and a re-derivation path. If memory grows,
514
- look for something bypassing a budget — not for a leak in the DAG (nodes are
515
- meant to accumulate).
516
-
517
- ### 2.13 Honest degradation, visible failure
518
-
519
- Nothing degrades silently: counters (`danglingReads`, `compactFailures`), trace
520
- steps (`bridgeMiss`, `narrowDecision`, `thinGrounding`, `skipMechanism` with the
521
- reason it skipped, `anchorFallback`), the `echoed` flag on recall's last tier,
522
- `recall-echo` provenance, and the per-region/per-anchor rejection reasons in the
523
- climb's structured payload. Empty results are legitimate outputs (silence), not
524
- errors — and several suites assert exactly that (5).
525
-
526
- _Follow it:_ when your code can degrade, emit a counter or trace step. And mind
527
- the classic trap: **empty bytes are truthy** — `Uint8Array(0)` passes
528
- `if (answer)`; always test `.length`.
529
-
530
- ### 2.14 Measured, not guessed — the work meter
531
-
532
- `src/meter.ts` is the one computational-usage accounting surface: a `Meter`
533
- counts the WORK one inference call performs at every layer (store reads by kind
534
- and by VOLUME, ANN queries and vectors scanned, perceptions, recognitions,
535
- climbs and ancestor visits, alignment cells, chart pops, mechanism floors/runs)
536
- and times named PHASES. Off by default and free when off;
537
- `new Mind({ profile:
538
- true })` attaches one per response and leaves a
539
- `CostReport` on `mind.lastCost`. `bench/profile-inference.mjs` is the reference
540
- harness.
541
-
542
- It is the profiling counterpart of the rationale — the rationale says why an
543
- answer was chosen, the meter says what it cost. Four contracts:
544
-
545
- 1. **Never read by inference.** A counter that reached a decision would end
546
- determinism (2.1). Write-only from the engine's side.
547
- 2. **Counts are the product; times are the hint.** Counters are deterministic,
548
- so two runs are diffable and a work regression is visible without a
549
- stopwatch. Only `elapsedMs` and the phase millisecond totals are not.
550
- 3. **Phases nest, and carry their own counter deltas.** `think` ⊃ `<mech>.run` ⊃
551
- `substitutionBridge`. Inclusive, never summed — but each phase reports the
552
- work done inside it (`PhaseCost.counters`), which is what makes "which phase
553
- did those byte reads?" answerable at all.
554
- 4. **Count a logical operation once.** A recursive read (`bytesPrefix`
555
- descending a branch) is charged at the public entry point only — the private
556
- `_prefix` body is uncharged. Counting the recursion made one read of an
557
- N-byte branch report as N reads, so the counter measured tree size instead of
558
- read requests.
559
- 5. **Shared analyses are charged to themselves.** `attention`, `weave`,
560
- `spanShaped` and `substitutionBridge` bill their own phase, not the mechanism
561
- that happened to first-touch them — otherwise the profile blames whoever paid
562
- on everyone's behalf (it once read "cast.floor costs 2.9 s" when 2.7 s of
563
- that was the consensus climb).
564
-
565
- _Follow it:_ a new layer that wants to be visible bumps a field in `meter.ts` —
566
- never a private counter. (`danglingReads`/`compactFailures` in `store.ts` stay:
567
- those are session-lifetime HEALTH counters, not per-response work.) Add the
568
- counter to the report the same way, and remember the classic trap: **profile
569
- without a trace attached** (2.11).
570
-
571
- ### 2.15 The fold contract: train and infer must agree
572
-
573
- `perceiveDeposit` and `perceive` must produce the SAME tree for the same bytes.
574
- That is not a nicety — it is what makes a trained context node and the node
575
- `resolve(query)` reaches the same node. The deposit path therefore imposes
576
- nothing: no boundaries, no turn convention, nothing read out of the bytes. When
577
- the two sides disagreed, the alignment family went quadratic (measured: 5.2M
578
- cells on a 476-byte context, against 0 when they agree) and cumulative contexts
579
- stopped resolving to what they were trained as.
580
-
581
- Three consequences you will meet:
582
-
583
- - **Boundaries are a separate feature from reuse.** `contentFoldIncremental`
584
- (segment reuse, transparent, imposes nothing) and `stablePrefixFold`
585
- (caller-supplied cuts, left-nested, buys prefix-ROOT identity) solve different
586
- problems. Conflating them is what once put an imposed boundary set on the
587
- inference path.
588
- - **A conversation's turn offsets are API metadata**, not a fold instruction.
589
- They feed `ConversationState`, `answeredSpans` and `currentTurnStart`; the
590
- geometry never sees them.
591
- - **Identity must not depend on W**, or on absolute offset. If you are adding
592
- anything that groups by index — a stride, a tile, a fixed-arity row — you are
593
- reintroducing the bug content-defined cuts exist to remove. Two such attempts
594
- are recorded as refuted at `collectRegions` and `contentLevels`; test/59 and
595
- test/63 pin the invariance floors.
596
-
597
- _Follow it:_ changes to `contentLevels` — the cut rate, which bits are read, the
598
- minimum/maximum segment length, the forced cut — are changes to the segment
599
- DISTRIBUTION every downstream mechanism is fitted to. Each of those four has
600
- been altered experimentally and cost 5–21 tests. Re-measure the whole suite, not
601
- the one query that motivated the change.
602
-
603
- ### 2.16 Comment style
604
-
605
- Comments state _constraints and failure modes_ — "this guard exists because X
606
- breaks without it", often naming the test that pins the behaviour — never
607
- narration of what the next line does. Two standing examples in
608
- `graph-search.ts`: the `hasHalo` guard on fusing completed rewrites (answer
609
- corruption via phrase-interior chunks) and the `couldGrow` liveness rule (O(N²)
610
- chart growth). When you fix a subtle bug, leave the constraint behind, not the
611
- story of the fix.
612
-
613
- ### 2.17 Saturation — every walk decides, none drifts
614
-
615
- Saturation is a first-class control, not a secondary nicety — and it is TWO
616
- things, distinguished deliberately:
617
-
618
- - a CAP — a derived bound (√N per read, √N·W per walk) that exists only to stop
619
- magic constants (§2.2) — is a SAFETY NET. It bounds the walk when its question
620
- never decides (a side too common to ever settle). Derived, never tuned; but it
621
- is not itself a decision.
622
- - a REAL saturation — a named, derived stop that DECIDES the walk's question and
623
- terminates the moment it is decided. The cap remains as the backstop; the
624
- saturation ends the walk. A walk with only a cap drifts to the cap every time;
625
- a walk with a real saturation stops where the answer is already known.
626
-
627
- `edgeAncestors` is the model (EXPAND-UNTIL-DECIDED): a reach is consumed either
628
- as a VOTE (needs `contextsReached` exactly, only while ≤ √N) or as an ABSTENTION
629
- (`saturated`), so it stops at the FIRST of its five named stops and no consumer
630
- reads a saturated reach's roots or counts. `pivotInto`'s candidate scan is the
631
- second: "longest valid wins" is DECIDED at the first valid candidate in
632
- descending length, so it reads one winner's bytes, never every shorter
633
- candidate. Saturation is a DECISION about the answer — never a cache, never a
634
- budget.
635
-
636
- The junction walk's per-node hub guards are real per-node saturations; its
637
- `√N·W` budget is the NET, not a saturation. REFUTED (test/16 bridge synthesis,
638
- test/34 n-ary binding): a "stop once one side's cone is exhausted" early stop is
639
- WRONG, in both a hub-guarded and a hub-flagged form. The junction test is a BYTE
640
- containment over the UNION of the two cones, and a junction can be structurally
641
- reachable from only ONE side — the side whose seed is a FOLD sub-node of the
642
- container. test/16: "cold or hot" is reached from the window "cold", but the
643
- 3-byte answer "hot" is not a 4-byte window of it, so "hot"'s cone empties after
644
- one pop while the junction still lies ahead in "cold"'s cone. "One cone
645
- exhausted" therefore never proves "no junction left", and the budget stays the
646
- net that backs the per-node saturations.
647
-
648
- _Follow it:_ when a new walk measures commonality against the corpus, name its
649
- saturation condition — the answer it may stop producing — beside its read cap.
650
- No code may traverse the inference uncontrolled from the corpus, nor lack the
651
- saturation its question admits. A cap without a saturation is a drift, and a
652
- drifting walk is a bug, not a tuning choice; do not mask it with a cache (§2.12)
653
- — a cache hides a drift on a warm store, saturation removes it.
654
-
655
- ---
58
+ Cross-cutting contracts (single-definition, import everywhere): `contentLevels`
59
+ in `src/geometry.ts` is the one boundary rule; `src/mind/canonical.ts` is the
60
+ write/read contract for canonical segmentation; `src/mind/junction.ts` is the
61
+ shared content-addressed ascent; `Precomputed` in
62
+ `src/mind/pipeline-mechanism.ts` is the per-response lazy memo; `src/meter.ts`
63
+ is the write-only work accounting surface. See `docs/INDEX.md` for the full
64
+ contract table and `docs/architecture/factored-machinery.md` for ownership.
656
65
 
657
66
  ## 3. Where things live
658
67
 
659
- | Concept | File(s) |
660
- | :-------------------------------------------------- | :------------------------------------------------------------------------------------------------- |
661
- | Public surface / assembly | `src/index.ts`, `src/mind/mind.ts` |
662
- | Conversation API (turns, state, answered spans) | `src/mind/mind.ts` |
663
- | Config (capacities, budgets, seed) | `src/config.ts` |
664
- | Derived thresholds, the fold, Hilbert | `src/geometry.ts` |
665
- | Content canonicalizer (injected, modality-specific) | `src/canon.ts` |
666
- | Vector primitives, alphabet, node/fold types, seats | `src/vec.ts`, `src/alphabet.ts`, `src/sema.ts` |
667
- | Perceive / resolve / read primitives | `src/mind/primitives.ts` |
668
- | Store domain logic / SQLite adapter | `src/store.ts`, `src/store-sqlite.ts` |
669
- | Mechanism contract + shared `Precomputed` | `src/mind/pipeline-mechanism.ts` |
670
- | The grounding decider (`think`) | `src/mind/pipeline.ts` |
671
- | Grounding mechanisms (one file each) | `src/mind/mechanisms/{cover,cast,confluence,extraction,reference,recall,prefix-completion,alu}.ts` |
672
- | Weighted deduction system + cost ladder | `src/mind/graph-search.ts` (engine in `src/derive/`) |
673
- | Match/project family | `src/mind/match.ts` |
674
- | Graph traversal, corpus scale, disambiguators | `src/mind/traverse.ts` |
675
- | Consensus climb + cross-region attention | `src/mind/attention.ts` |
676
- | Recall's refusal-path tier (substitution bridge) | `src/mind/bridge.ts` |
677
- | Recognition / canonical contract | `src/mind/recognition.ts`, `src/mind/canonical.ts` |
678
- | Junction ascent (bridge + attention share) | `src/mind/junction.ts`, `src/mind/resonance.ts` |
679
- | Learning / ingestion / training cache | `src/mind/learning.ts`, `src/ingest-cache.ts` |
680
- | Post-grounding (reason, fuse, articulate) | `src/mind/reasoning.ts`, `src/mind/articulation.ts` |
681
- | Rationale / trace | `src/mind/rationale.ts`, `src/mind/trace.ts` |
682
- | Computational-usage meter | `src/meter.ts` (harness: `bench/profile-inference.mjs`) |
683
- | Extension host types | `src/extension.ts` |
684
- | Sublibraries (own READMEs, own tests) | `src/derive/`, `src/alu/`, `src/rabitq-ivf/` |
685
-
686
- Mind functions are **free functions over `MindContext`** (`mind/types.ts`), not
687
- methods — `mind.ts` is a thin assembly that implements the context and
688
- delegates. Follow that shape: it keeps every mechanism testable in isolation
689
- with no hidden `this` state.
690
-
691
- ---
68
+ | Concept | File(s) |
69
+ | ------------------------------------ | -------------------------------------------------------------------------------------------------- |
70
+ | Public surface / assembly | `src/index.ts`, `src/mind/mind.ts` |
71
+ | Config (capacities, budgets, seed) | `src/config.ts` |
72
+ | Derived thresholds, fold, Hilbert | `src/geometry.ts` |
73
+ | Content canonicalizer (injected) | `src/canon.ts` |
74
+ | Vector primitives, alphabet, seats | `src/vec.ts`, `src/alphabet.ts`, `src/sema.ts` |
75
+ | Perceive / resolve / read primitives | `src/mind/primitives.ts` |
76
+ | Store domain logic / SQLite adapter | `src/store.ts`, `src/store-sqlite.ts` |
77
+ | Mechanism contract + `Precomputed` | `src/mind/pipeline-mechanism.ts` |
78
+ | Grounding decider (`think`) | `src/mind/pipeline.ts` |
79
+ | Grounding mechanisms (one file each) | `src/mind/mechanisms/{cover,cast,confluence,extraction,reference,recall,prefix-completion,alu}.ts` |
80
+ | Weighted deduction + cost ladder | `src/mind/graph-search.ts` (engine in `src/derive/`) |
81
+ | Match/project family | `src/mind/match.ts` |
82
+ | Graph traversal, corpus scale | `src/mind/traverse.ts` |
83
+ | Consensus climb + attention | `src/mind/attention.ts` |
84
+ | Substitution bridge (recall tier) | `src/mind/bridge.ts` |
85
+ | Recognition / junction / resonance | `src/mind/recognition.ts`, `src/mind/junction.ts`, `src/mind/resonance.ts` |
86
+ | Learning / ingestion | `src/mind/learning.ts`, `src/ingest-cache.ts` |
87
+ | Rationale / trace / meter | `src/mind/rationale.ts`, `src/mind/trace.ts`, `src/meter.ts` |
88
+ | Extension host types | `src/extension.ts` |
89
+ | Sublibraries (own READMEs) | `src/derive/`, `src/alu/`, `src/rabitq-ivf/` |
90
+
91
+ Mind functions are free functions over `MindContext` (`src/mind/types.ts`), not
92
+ methods — `mind.ts` is a thin assembly that delegates.
692
93
 
693
94
  ## 4. Recipes
694
95
 
695
96
  ### Add a grounding mechanism or extension
696
97
 
697
- Implement `PipelineMechanism`: `floor` returns an admissible bound or `null`
698
- (structurally can't fire); `run` returns candidates with `bytes`, `accounted`,
699
- `moves`, `unexplained` (plus `scaffolding`/`complete` when your mechanism can
700
- state them); add `parse` only if you compute authoritative spans (the ALU's
701
- `aluToMechanism` in `mechanisms/alu.ts` is the reference). Register with
702
- `new Mind({ mechanismFactories: [host => yourMechanism(host)] })` (or
703
- `mechanisms: [...]` if no host is needed); reach meaning only through the
704
- `ExtensionHost`. Verify the four market constraints (2.6). You never touch
705
- `think()` or another mechanism's file.
706
-
707
- If your `floor` needs an expensive shared analysis to be tight, check
708
- `worthRunning(cheapestBound)` FIRST and return the uninvested bound when it
709
- already fails — that is the investment discipline (2.6), and `cast.ts` /
710
- `extraction.ts` are the two reference implementations.
98
+ Implement `PipelineMechanism` (`floor` → admissible bound or `null`; `run` →
99
+ candidates with `bytes`/`accounted`/`moves`/`unexplained` + optional
100
+ `scaffolding`/`complete`). Register via
101
+ `new Mind({ mechanismFactories: [host => yourMechanism(host)] })`. Verify the
102
+ four market constraints (decoupled, declared competence, visible budget,
103
+ evidence travels). → `docs/architecture/mechanism-market.md`
711
104
 
712
105
  ### Add an ALU operation
713
106
 
714
- One declarative `registry.derive(name, arity, surfaceForms, body)` in the
715
- relevant `src/alu/src/kernel-*.ts`; the body composes existing ops. Scalar ops
716
- broadcast over n-d automatically. No parser, search, or mind edits — see
717
- `src/alu/README.md`.
107
+ One `registry.derive(name, arity, surfaceForms, body)` in
108
+ `src/alu/src/kernel-*.ts` composing existing ops. No parser/search/mind edits. →
109
+ `src/alu/README.md`
718
110
 
719
111
  ### Add a deduction rule
720
112
 
721
113
  Rules live in `GraphSearch` (`coverRules`/`formRules`/`outRules`/`fuse`). Place
722
- its cost in the ladder deliberately (2.4), emit it lazily from the item kind
723
- that triggers it, keep the heuristic admissible, extend `classifyMove` (the
724
- single rule-shape → move-name mapping for the rationale), and add a
725
- rationale-visible test. Async data is pre-resolved in the pipeline (2.10).
114
+ cost deliberately on the ladder, keep the A* heuristic admissible, extend
115
+ `classifyMove`, add a rationale-visible test. Pre-resolve async data in the
116
+ pipeline. → `docs/architecture/cost-model.md`
726
117
 
727
118
  ### Add a store backend
728
119
 
729
- Subclass `AbstractStore`; implement the `_db*`/`_vec*` methods as thin statement
730
- wrappers (`SQliteStore` is the template); LIMITed variants must be real `LIMIT`
731
- queries and existence probes real point probes (2.8, 2.9). Run the full suite
732
- with your store substituted.
733
-
734
- ### Add a modality
735
-
736
- Perception consumes byte streams. Grid-shaped data: build a `Grid`
737
- (`{width, height, channels, data}` or n-dimensional `dims`) — `geometry.ts`
738
- Hilbert-linearizes it; `Grid[]` stacks frames. Anything else: produce a
739
- `Uint8Array` with a deterministic, locality-preserving ordering. Nothing
740
- downstream changes.
741
-
742
- A modality may also supply its own **canonicalizer** (`Canon`) and its own
743
- reading of "edge" — that is the one place presentation rules belong. Nothing in
744
- the store or the mind's core knows what case or whitespace is; the text entry
745
- points inject `textCanon`/`textEdgeTrim`, byte and grid inputs inject neither
746
- (for them `0x20` is content). If you find yourself adding a character class
747
- inside a mechanism, it belongs here instead.
748
-
749
- ### Hold a conversation
750
-
751
- ```ts
752
- const conv = mind.beginConversation(savedState); // state optional
753
- const { response, state } = await mind.respondTurnText(conv, "…");
754
- mind.addTurn(conv, "…"); // a turn to hear but not answer
755
- mind.endConversation(conv);
756
- ```
757
-
758
- Turns append raw bytes plus an offset — never a separator. To replay a corpus
759
- that joins turns with `"\n"`, pass `"\n" + turnText` as the turn; the separator
760
- rides inside the turn bytes, where it belongs. `ConversationState` (context,
761
- boundaries, answered spans) is serialisable and restores exactly. One
762
- `respondTurn` may be in flight per Mind: the conversation's memos are swapped
763
- into the response-scoped slots for the turn's duration.
764
-
765
- ### Debug an answer
766
-
767
- ```ts
768
- const r = await mind.respond(query, (rationale) => {
769
- console.dir(rationale, { depth: null }); // every step, cost, data-flow edge
770
- });
771
- console.log(r.provenance); // cast | join | cover | extract | reference | recall | recall-echo | prefix
772
- ```
773
-
774
- Read top-down: which mechanism fired (and why the others abstained), what
775
- recognition found, how the climb voted, which edges were followed
776
- (`disambiguate` steps carry the evidence). `recall-echo` means "nearest stored
777
- form, not a derived fact". `reference` means "part of this answer is bytes the
778
- ASKER supplied, voiced through a slot the corpus attests as a carriage" — read
779
- its `bindReferent` step for the referents and the instances, and
780
- `referenceLicence` for why a binding was refused. `prefix` means "the query is
781
- the literal opening of exactly one trained form, which this voiced whole".
782
-
783
- Three steps carry **structured data**, so tooling need not parse notes:
784
- `decideGrounding` (every candidate's provenance, exact weight, discrete grade,
785
- unexplained bytes, which won, plus the runner-up margin), `climbConsensus` (per
786
- region: source, span, selected anchor, IDF, contrastive margin and its rival,
787
- mutual weight, outcome; per anchor: pooled vote, peak, breadth, clusters, and
788
- the live commit verdict with its rejection reasons; plus every cross-region
789
- probe and its tier), and `narrowDecision`. When an answer is wrong, the fastest
790
- route is usually `decideGrounding` first (was the right mechanism outbid, or did
791
- it never produce a candidate?), then the per-region climb detail (did the
792
- evidence vote at all?).
793
-
794
- ### Profile an answer
795
-
796
- ```ts
797
- const mind = new Mind({ store, profile: true }); // no trace — see 2.14
798
- await mind.respondText(query);
799
- console.log(formatReport(mind.lastCost!)); // counters by layer + nested phases
800
- ```
801
-
802
- `bench/profile-inference.mjs` runs a battery this way and prints per-query and
803
- aggregate reports; `sumReports` aggregates a run. Diff the COUNTERS between two
804
- runs to catch a work regression — they are deterministic; read the PHASES to see
805
- where the wall clock went.
806
-
807
- ### Train at scale
808
-
809
- Use `CachedIngest` (`ingest-cache.ts`) as a drop-in for `mind.ingest` — it
810
- memoises perceive+intern of repeated inputs and routes through the same
811
- `dispatchIngest` as the direct path (shape detection can't drift). Call
812
- `store.commit()` at checkpoints; run `compactContentIndex` /
813
- `repairContentIndex` post-training if eviction was heavy, and
814
- `mind.buildCanonIndex()` if queries will carry a canonicalizer (2.9). See
815
- `example/train_base/` — a folder, entry `main.ts`: the run context lives in
816
- `runtime.ts`, the one per-corpus loop in `stage.ts`, and each corpus (knobs, row
817
- adapter, stage descriptor) in `corpora/<name>.ts`. Profiling note: the first
818
- `resonate` after a big ingest pays the pending index flush; the dominant
819
- query-side ANN cost is connector pre-resolution (bounded by recognised-site
820
- count — don't add another loop over site pairs).
821
-
822
- ---
120
+ Subclass `AbstractStore`; implement `_db*`/`_vec*` as thin wrappers. `LIMIT`
121
+ variants must be real `LIMIT ?` queries; existence probes must be point probes.
122
+ Run the full suite with your store substituted. → `docs/architecture/store.md`
123
+ (+ `docs/architecture/bounded-reads.md` for caps)
823
124
 
824
125
  ## 5. Testing norms
825
126
 
826
- Tests are plain `node:test` suites in `test/*.test.mjs`, numbered by theme, run
827
- against the built `dist/` (`npm test`; one suite:
828
- `node --test test/22-multihop.test.mjs` after `tsc`).
829
-
830
- - New behaviour ⇒ a test in the matching numbered suite, or a new numbered
831
- suite.
832
- - Many tests pin **contracts that look like implementation details** (the bridge
833
- tier order, the bridge's identity admission vs. its prefix trap and the
834
- scaffolding reading behind it, the two span-shape readings (`match.ts`),
835
- `MechanismResult.complete`, fold invariance under shifts (test/59, test/63),
836
- recognition's idempotence under trace (test/42), the cross-region tier ladder
837
- (test/51), the instrumentation payloads (test/52–55), determinism, honest
838
- silence). A "simplification" that fails an existing test is wrong until you
839
- can argue the _test_ is wrong — several guards exist precisely because a
840
- plausible simplification once failed a dozen suites.
841
- - **Honest silence is a tested behaviour, not an absence of one.** Several
842
- suites assert that a query grounds NOTHING; a change that makes the engine
843
- more forthcoming fails them, and that is the suite working.
844
- - Sublibraries test themselves (`src/{alu,derive,rabitq-ivf}/test/`) with zero
845
- Sema dependency. Keep it so.
846
- - Performance claims are tested (the rabitq-ivf benchmark asserts sub-linear
847
- scaling and compression). Changing index behaviour means running it.
848
-
849
- ---
127
+ Tests are `node:test` suites in `test/*.test.mjs`, numbered by theme, run
128
+ against built `dist/` (`npm test`; single suite:
129
+ `node --test test/22-multihop.test.mjs` after `tsc`). New behaviour ⇒ test in
130
+ the matching numbered suite or a new one. Many tests pin contracts that look
131
+ like implementation details (ladder order, span-shape readings,
132
+ `MechanismResult.complete`, fold invariance, recognition idempotence, honest
133
+ silence). A simplification that fails an existing test is wrong until the test
134
+ is proven wrong. Sublibraries test themselves in
135
+ `src/{alu,derive,rabitq-ivf}/test/` with zero Sema dependency.
850
136
 
851
137
  ## 6. Dependencies and licensing
852
138
 
853
139
  PolyForm Noncommercial 1.0.0 with separate commercial licensing (see
854
- [LICENSE.md](LICENSE.md), [COMMERCIAL-LICENSE.md](COMMERCIAL-LICENSE.md),
855
- [TRADEMARKS.md](TRADEMARKS.md), [CONTRIBUTING.md](CONTRIBUTING.md)). Do not
856
- vendor code under licenses incompatible with dual distribution, and do not add
857
- runtime dependencies casually — the near-zero-dependency footprint is a product
858
- feature.
859
-
860
- **The library has NO runtime dependencies at all**, and that is now pinned by
861
- `test/88-dependency-footprint.test.mjs`: the built `dist/src` may import only
862
- `node:` builtins and relative paths, `package.json` may declare no
863
- `dependencies`, and the published entry points may not reach outside `dist/src`.
864
- The rule an EXAMPLE follows is different and looser — it may use what it needs,
865
- as a **dev** dependency, loaded **lazily** so it is a requirement only of the
866
- code path that uses it. `example/train_base` is the reference: `hyparquet` (+
867
- its Snappy codec) is the sole third-party package in this repository, it is
868
- dev-only, and `readers.ts` resolves it by dynamic import the first time a
869
- Parquet corpus is actually read — so a curriculum with no Parquet stage runs
870
- with the package absent. It used to sit in `dependencies`, which installed a
871
- Parquet reader on every consumer of Sema for the sake of one example; that is
872
- the mistake the suite exists to catch.
873
-
874
- **Training corpora are governed by the same rule, and more strictly.** Sema is
875
- non-parametric: a trained store retains its training text VERBATIM (read any
876
- content node back and the original sentence comes out). A store is therefore a
877
- redistribution of its corpora, not a derived model, and every upstream licence
878
- applies to it in full. Two consequences:
879
-
880
- - A corpus carrying a **NonCommercial** term cannot enter a trainer — it
881
- conflicts with the commercial licence tier.
882
- - A corpus carrying a **ShareAlike** term cannot enter a trainer — its copyleft
883
- would attach to the distributed store.
884
-
885
- Check both against **what the corpus was built from**, not only the repository's
886
- license tag: a dataset assembled out of Wikipedia prose and published under
887
- Apache-2.0 still carries CC BY-SA on that prose. Where a corpus has a clean
888
- layer and a contaminated one, ingest only the clean layer.
889
-
890
- Sema's own license does **not** extend over corpus content inside a store, and
891
- cannot: CC BY 4.0 §2(a)(5)(B) forbids applying terms that restrict what the
892
- license permits. The engine is what PolyForm protects. Per-corpus attribution,
893
- the required modification statement, and the current allow/deny list live in
894
- [DATASETS.md](DATASETS.md) — update it in the same change that touches a
895
- trainer's corpus set.
140
+ `LICENSE.md`, `COMMERCIAL-LICENSE.md`, `TRADEMARKS.md`). The library has **no
141
+ runtime dependencies** — pinned by `test/88-dependency-footprint.test.mjs`
142
+ (`dist/src` may import only `node:` builtins and relative paths; `package.json`
143
+ has no `dependencies`). Examples may use dev dependencies lazily via dynamic
144
+ import only on the code path that needs them (`example/train_base` + `hyparquet`
145
+ is the reference). Training corpora: a store retains text verbatim, so upstream
146
+ licences apply in full — NonCommercial and ShareAlike corpora cannot enter a
147
+ trainer; see `DATASETS.md`.